K-Means Clustering
Unsupervised Learning Concepts
Unlike supervised learning, unsupervised learning works with data that has no labels. The goal is to discover hidden patterns, structures, or groupings in the data.
Types of Unsupervised Learning
Why Clustering Matters
Applications:
- Customer segmentation
- Document topic grouping
- Image segmentation
- Anomaly detection
- Gene expression analysis
- Social network analysis
K-Means Algorithm
K-Means is one of the most popular and straightforward clustering algorithms.
Algorithm Steps
Mathematical Formulation
Complete Python Implementation
Finding Optimal K
Elbow Method
Plot inertia vs K and look for the "elbow" – the point where adding more clusters provides diminishing returns.
Silhouette Score
- : Well-clustered
- : On cluster boundary
- : Likely in wrong cluster
Architecture Diagram
Elbow plot + Silhouette plot:
K=2: Inertia high, Silhouette moderate
K=3: Inertia dropping, Silhouette high
K=4: Inertia elbow, Silhouette highest <- Optimal
K=5+: Inertia plateaus, Silhouette decreasing
K-Means Variants
Mini-Batch K-Means
For large datasets, Mini-Batch K-Means uses random subsets for faster training:
Limitations and Alternatives
K-Means Limitations
Alternatives Comparison
Algorithm Comparison Summary
Practical Application
Key Takeaways
Practice Exercises
Exercise 1: Image Segmentation
Exercise 2: Anomaly Detection
Exercise 3: Compare Clustering Algorithms
Exercise 4: Choosing K for Real Data
Summary Table
| Algorithm | K Required | Shape Assumption | Complexity | Best For |
|---|---|---|---|---|
| K-Means | Yes | Spherical | O(nK) | Spherical clusters |
| DBSCAN | No | Arbitrary | O(n log n) | Noisy, irregular shapes |
| Hierarchical | No | Any | O(n²) | Small datasets |
| GMM | Yes | Elliptical | O(nK) | Soft assignments |