sciBASIC# knot logo sciBASIC# ↖

04 Tutorial — K-means + PCA

Cluster the Iris dataset with k-means, then unfold it onto a 2-D PCA scatter.

Module DataMining · KMeans + PCA Dataset bezdekIris · 150 × 4 Clusters k = 3 Output PNG + CSV

A complete unsupervised-learning pass over the classic Bezdek Iris table — 150 specimens, four morphological measurements, three hidden species — driven by a short VBScript run on the sciBASIC# script engine. The script loads the unified NumericTable straight from CSV, partitions it into k = 3 clusters, projects the four-dimensional points onto the first two principal components, renders an 800 × 600 Nature-themed scatter, and exports the fully annotated table as CSV.

02 Pipeline

One table, end to end

Step 1

Load

NumericTableIO.ReadCsv reads the CSV into the unified NumericTable — the columns whitelist keeps only the four feature columns D1–D4; the sample-id column and the text species column are ignored.

Step 2

Cluster

table.kmeans(k := 3) partitions the 150 specimens and writes the result back onto the very same table as the label:cluster column.

Step 3

Reduce

result.pca(maxPC := 2).ScoreTable() unfolds the 4-D points onto the PC1 / PC2 plane as a new features table.

Step 4

Plot

ClusterLabels() recovers the cluster colors; a ScatterPlot (800 × 600, PlotTheme.Nature) draws PC1 vs PC2 colored by cluster.

Step 5

Export

plt.SavePng writes the 300-dpi figure; result.WriteCsv streams the annotated table back to CSV in the same label: layout, so NumericTableIO.ReadCsv can reload it losslessly.

01 The Script

Full demo source

The complete script exactly as executed by the sciBASIC# script engine (vbs.exe) — nothing elided.

kmeans.vb · 88 linesDownload kmeans.vb
#include "Microsoft.VisualBasic.DataMining.Framework.dll"
#include "Microsoft.VisualBasic.Data.Framework.dll"
#include "Microsoft.VisualBasic.Data.DataPlot.dll"
#include "Microsoft.VisualBasic.Math.Statistics.ANOVA.dll"
#include "Microsoft.VisualBasic.Drawing.dll"

imports microsoft.visualbasic.data
imports microsoft.visualbasic.data.framework
imports microsoft.visualbasic.datamining.kmeans
imports microsoft.visualbasic.datamining
imports Microsoft.VisualBasic.Math.Statistics.Hypothesis.ANOVA
imports microsoft.visualbasic.data.plots
imports microsoft.visualbasic.drawing

dim file = here("../../data/bezdekIris.csv")
dim k = 3

' ---------------------------------------------------------------------------
' 1. Load the unified 2D table object (NumericTable) directly from the csv file
'
'    The table layout convention is "row names + feature columns + label
'    columns prefixed by label:", so the columns whitelist below loads only the
'    four numeric feature columns D1..D4; the first column of the csv file
'    (the sample id) and the text class column are both ignored
' ---------------------------------------------------------------------------
dim table = NumericTableIO.ReadCsv(file, columns := {"D1","D2","D3","D4"})

call console.WriteLine($"load table: {table.nsamples} samples x {table.nfeatures} features")

' ---------------------------------------------------------------------------
' 2. Run the KMeans clustering
'
'    The clustering result is written into the label matrix of the table as
'    the label column cluster, so the following analysis steps can reuse the
'    very same table
' ---------------------------------------------------------------------------
dim result = table.kmeans(k := k)

' ---------------------------------------------------------------------------
' 3. Run PCA on the clustered result table
'
'    The pca method accepts the unified 2D table and its result is still a
'    MultivariateAnalysisResult object; the ScoreTable extension method then
'    extracts the projected coordinates (scores) into a new table:
'
'    + rows     = samples
'    + features = the PC1, PC2 principal component scores
' ---------------------------------------------------------------------------
dim score = result.pca(maxPC := 2).ScoreTable()

' ---------------------------------------------------------------------------
' 4. Extract the data needed for plotting from the result table
' ---------------------------------------------------------------------------
dim pc1 = score.Feature("PC1")
dim pc2 = score.Feature("PC2")
dim classes = result.ClusterLabels()
dim class_id(classes.length - 1) as string
dim sizes(k - 1) as integer

for i = 0 to classes.length - 1
    class_id(i) = classes(i).ToString()
    sizes(classes(i) - 1) += 1
next

call console.WriteLine($"kmeans cluster sizes: {String.Join(", ", sizes)}")

Call SkiaDriver.Register()

Using plt As New ScatterPlot(800, 600, PlotTheme.Nature())
    plt.Title = "PCA group of bezdek-Iris"
    plt.SubTitle = "PCA score scatter with 3 iris species colors"
    plt.XLabel = "PC1"
    plt.YLabel = "PC2"
    plt.Plot(DataSerials(x := pc1, y := pc2, class_id).tolist())
    plt.SavePng(here("bezdekIris-pca-groups.png"), 300)
End Using

' ---------------------------------------------------------------------------
' 5. Export the clustering result table as a csv file
'
'    The exported file layout still follows the convention of
'    "row names + feature columns + label: prefixed label columns", so it can
'    be loaded back losslessly through NumericTableIO.ReadCsv
' ---------------------------------------------------------------------------
call result.WriteCsv(here("bezdekIris-kmeans.csv"))

call console.WriteLine("done: bezdekIris-pca-groups.png")
call console.WriteLine("done: bezdekIris-kmeans.csv")

03 Results

PCA scatter & cluster table

PCA score scatter of the bezdekIris dataset, 150 points colored by k-means cluster
Fig. 1 — PC1 / PC2 score scatter of the Bezdek Iris table, 150 specimens colored by k-means cluster. Cluster 1 (red) — I. setosa — occupies a region of its own on the left; clusters 2 and 3 (blue / green) partition the morphologically intergrading versicolor–virginica range.

Table preview — bezdekIris-kmeans.csv

The exported table carries the assigned cluster number (label:cluster) next to the four raw measurements, in the standard row names + feature columns + label: columns layout. First rows, a middle slice and the last rows are shown; the full file holds all 150 data rows.

D1D2D3D4label:cluster
Iris-setosa5.13.51.40.21
Iris-setosa4.931.40.21
Iris-setosa4.73.21.30.21
Iris-setosa4.63.11.50.21
Iris-setosa53.61.40.21
···
Iris-versicolor73.24.71.42
Iris-versicolor6.43.24.51.52
Iris-versicolor6.93.14.91.53
···
Iris-virginica6.535.223
Iris-virginica6.23.45.42.33
Iris-virginica5.935.11.82
Run summaryValue
Load table150 samples × 4 features
PCA components2 (17 + 10 loops)
Cluster 150
Cluster 262
Cluster 338
K-means recovers I. setosa perfectly — the 50 specimens of cluster 1 all sit inside their own region of the PC plane — while the overlapping versicolor / virginica pair is split 62 vs 38 across clusters 2 and 3, the textbook result for this dataset.