04 Tutorial — K-means + PCA
Cluster the Iris dataset with k-means, then unfold it onto a 2-D PCA scatter.
A complete unsupervised-learning pass over the classic Bezdek Iris table — 150 specimens, four morphological measurements, three hidden species — driven by a short VBScript run on the sciBASIC# script engine. The script loads the unified NumericTable straight from CSV, partitions it into k = 3 clusters, projects the four-dimensional points onto the first two principal components, renders an 800 × 600 Nature-themed scatter, and exports the fully annotated table as CSV.
02 Pipeline
One table, end to end
Load
NumericTableIO.ReadCsv reads the CSV into the unified NumericTable — the
columns whitelist keeps only the four feature columns D1–D4; the sample-id column and the
text species column are ignored.
Cluster
table.kmeans(k := 3) partitions the 150 specimens and writes the result back onto the very
same table as the label:cluster column.
Reduce
result.pca(maxPC := 2).ScoreTable() unfolds the 4-D points onto the PC1 / PC2 plane as a
new features table.
Plot
ClusterLabels() recovers the cluster colors; a ScatterPlot (800 × 600,
PlotTheme.Nature) draws PC1 vs PC2 colored by cluster.
Export
plt.SavePng writes the 300-dpi figure; result.WriteCsv streams the annotated
table back to CSV in the same label: layout, so NumericTableIO.ReadCsv can reload
it losslessly.
01 The Script
Full demo source
The complete script exactly as executed by the sciBASIC# script engine (vbs.exe) — nothing elided.
#include "Microsoft.VisualBasic.DataMining.Framework.dll"
#include "Microsoft.VisualBasic.Data.Framework.dll"
#include "Microsoft.VisualBasic.Data.DataPlot.dll"
#include "Microsoft.VisualBasic.Math.Statistics.ANOVA.dll"
#include "Microsoft.VisualBasic.Drawing.dll"
imports microsoft.visualbasic.data
imports microsoft.visualbasic.data.framework
imports microsoft.visualbasic.datamining.kmeans
imports microsoft.visualbasic.datamining
imports Microsoft.VisualBasic.Math.Statistics.Hypothesis.ANOVA
imports microsoft.visualbasic.data.plots
imports microsoft.visualbasic.drawing
dim file = here("../../data/bezdekIris.csv")
dim k = 3
' ---------------------------------------------------------------------------
' 1. Load the unified 2D table object (NumericTable) directly from the csv file
'
' The table layout convention is "row names + feature columns + label
' columns prefixed by label:", so the columns whitelist below loads only the
' four numeric feature columns D1..D4; the first column of the csv file
' (the sample id) and the text class column are both ignored
' ---------------------------------------------------------------------------
dim table = NumericTableIO.ReadCsv(file, columns := {"D1","D2","D3","D4"})
call console.WriteLine($"load table: {table.nsamples} samples x {table.nfeatures} features")
' ---------------------------------------------------------------------------
' 2. Run the KMeans clustering
'
' The clustering result is written into the label matrix of the table as
' the label column cluster, so the following analysis steps can reuse the
' very same table
' ---------------------------------------------------------------------------
dim result = table.kmeans(k := k)
' ---------------------------------------------------------------------------
' 3. Run PCA on the clustered result table
'
' The pca method accepts the unified 2D table and its result is still a
' MultivariateAnalysisResult object; the ScoreTable extension method then
' extracts the projected coordinates (scores) into a new table:
'
' + rows = samples
' + features = the PC1, PC2 principal component scores
' ---------------------------------------------------------------------------
dim score = result.pca(maxPC := 2).ScoreTable()
' ---------------------------------------------------------------------------
' 4. Extract the data needed for plotting from the result table
' ---------------------------------------------------------------------------
dim pc1 = score.Feature("PC1")
dim pc2 = score.Feature("PC2")
dim classes = result.ClusterLabels()
dim class_id(classes.length - 1) as string
dim sizes(k - 1) as integer
for i = 0 to classes.length - 1
class_id(i) = classes(i).ToString()
sizes(classes(i) - 1) += 1
next
call console.WriteLine($"kmeans cluster sizes: {String.Join(", ", sizes)}")
Call SkiaDriver.Register()
Using plt As New ScatterPlot(800, 600, PlotTheme.Nature())
plt.Title = "PCA group of bezdek-Iris"
plt.SubTitle = "PCA score scatter with 3 iris species colors"
plt.XLabel = "PC1"
plt.YLabel = "PC2"
plt.Plot(DataSerials(x := pc1, y := pc2, class_id).tolist())
plt.SavePng(here("bezdekIris-pca-groups.png"), 300)
End Using
' ---------------------------------------------------------------------------
' 5. Export the clustering result table as a csv file
'
' The exported file layout still follows the convention of
' "row names + feature columns + label: prefixed label columns", so it can
' be loaded back losslessly through NumericTableIO.ReadCsv
' ---------------------------------------------------------------------------
call result.WriteCsv(here("bezdekIris-kmeans.csv"))
call console.WriteLine("done: bezdekIris-pca-groups.png")
call console.WriteLine("done: bezdekIris-kmeans.csv")
03 Results
PCA scatter & cluster table
Table preview — bezdekIris-kmeans.csv
The exported table carries the assigned cluster number (label:cluster) next to the
four raw measurements, in the standard row names + feature columns + label: columns
layout. First rows, a middle slice and the last rows are shown; the full file holds all 150 data rows.
| D1 | D2 | D3 | D4 | label:cluster | |
|---|---|---|---|---|---|
| Iris-setosa | 5.1 | 3.5 | 1.4 | 0.2 | 1 |
| Iris-setosa | 4.9 | 3 | 1.4 | 0.2 | 1 |
| Iris-setosa | 4.7 | 3.2 | 1.3 | 0.2 | 1 |
| Iris-setosa | 4.6 | 3.1 | 1.5 | 0.2 | 1 |
| Iris-setosa | 5 | 3.6 | 1.4 | 0.2 | 1 |
| ··· | |||||
| Iris-versicolor | 7 | 3.2 | 4.7 | 1.4 | 2 |
| Iris-versicolor | 6.4 | 3.2 | 4.5 | 1.5 | 2 |
| Iris-versicolor | 6.9 | 3.1 | 4.9 | 1.5 | 3 |
| ··· | |||||
| Iris-virginica | 6.5 | 3 | 5.2 | 2 | 3 |
| Iris-virginica | 6.2 | 3.4 | 5.4 | 2.3 | 3 |
| Iris-virginica | 5.9 | 3 | 5.1 | 1.8 | 2 |
| Run summary | Value |
|---|---|
| Load table | 150 samples × 4 features |
| PCA components | 2 (17 + 10 loops) |
| Cluster 1 | 50 |
| Cluster 2 | 62 |
| Cluster 3 | 38 |