sciBASIC# knot logo sciBASIC# ↖

06 Tutorial — Word embeddings

Train CBOW word vectors on the tale of Rapunzel, then unfold the embedding with UMAP and k-means.

Module Data.NLP · Word2Vec Corpus Rapunzel.txt · CBOW Vectors 461 words · 30-D Embedding UMAP 9-D · 64 neighbors · k = 9

One fairy tale, one model, one picture of a vocabulary. The Grimm text of Rapunzel is segmented into sentences and fed to a continuous bag-of-words trainer (30 dimensions, min term frequency 1, single thread). The resulting 461 word vectors are projected onto a 9-dimensional UMAP manifold, partitioned into nine k-means clusters, and drawn as a Nature-themed scatter of UMAP1 vs UMAP2 — with the full coordinates exported to CSV.

02 Pipeline

From fairy tale to 9-D embedding

Step 1

Segment

Paragraph.Segmentation splits the Rapunzel text into paragraphs of sentences, and every sentence stream is handed to wv.readTokens.

Step 2

Train

BuildWord2VecFactory configures a continuous bag-of-words model — setVectorSize(30), TrainMethod.CBow, setFreqThresold(1), a single thread — and wv.training() runs the pass over the tokens.

Step 3

Embed

The 461 word vectors are pushed through umap(dimensions:=9, numberOfNeighbors:=64): InitializeFit + Step(n_epochs) produce the 9-D manifold.

Step 4

Cluster

Kmeans(k:=9) partitions the embedding; ClassId recovers cluster colors and the first two UMAP dimensions become the plot coordinates.

Step 5

Export

A Nature-themed scatter saves the 300-dpi PNG, and the full 461 × 11 table is written to CSV with WriteCsv.

01 The Script

Full demo source

The complete script exactly as executed by the sciBASIC# script engine (vbs.exe) — nothing elided.

word2vector.vb · 129 linesDownload word2vector.vb
#include "Microsoft.VisualBasic.Data.NLP.Word2Vec.dll"
#include "Microsoft.VisualBasic.Data.NLP.dll"
#include "Microsoft.VisualBasic.DataMining.Framework.dll"
#include "Microsoft.VisualBasic.Drawing.dll"
#include "Microsoft.VisualBasic.Data.DataPlot.dll"
#include "Microsoft.VisualBasic.Data.Framework.dll"
#include "Microsoft.VisualBasic.Math.Randomizer.dll"
#include "Microsoft.VisualBasic.DataMining.UMAP.dll"

Imports Microsoft.VisualBasic.Data.NLP.Word2Vec
Imports Microsoft.VisualBasic.Data.NLP.Model
Imports Microsoft.VisualBasic.Data
Imports Microsoft.VisualBasic.Data.Framework
Imports Microsoft.VisualBasic.DataMining
Imports Microsoft.VisualBasic.DataMining.Kmeans
Imports Microsoft.VisualBasic.DataMining.UMAP
imports microsoft.visualbasic.data.plots
imports microsoft.visualbasic.drawing

' ---------------------------------------------------------------------------
' Word2Vec + UMAP + KMeans demo
'
'   Rapunzel text -> Word2Vec training -> word vector table
'    -> UMAP embedding -> KMeans clustering -> scatter plot -> export csv
' ---------------------------------------------------------------------------
dim textfile = here("../../data/Rapunzel.txt")
dim wv As Word2Vec = BuildWord2VecFactory() _
    .setVectorSize(30) _
    .setMethod(TrainMethod.CBow) _
    .setNumOfThread(1) _
    .setFreqThresold(1) _
    .build()
Dim data As Paragraph() = Paragraph.Segmentation(textFile.ReadAllText).ToArray

For Each p As Paragraph In data
    For Each line In p.sentences
        Call wv.readTokens(line)
    Next
Next

call wv.training()

' ---------------------------------------------------------------------------
' 1. Build the trained word vector collection into a unified 2D table object
'    (NumericTable)
'
'    + row names  = the word token
'    + features   = v_1 .. v_n, i.e. the word vector of each token
'
'    All the following operations (dimension reduction, clustering, export) are
'    performed directly on this table
' ---------------------------------------------------------------------------
dim vector = wv.outputVector()
dim tokens = vector.tokens
dim colnames = fieldname("v", n := vector.vectorSize).toarray()
dim rows As Double()() = New Double(tokens.length - 1)() {}

for i = 0 to tokens.length - 1
    dim raw = vector.wordMap(tokens(i))
    dim vals(raw.length - 1) as double

    for j = 0 to raw.length - 1
        vals(j) = CDbl(raw(j))
    next

    rows(i) = vals
next

dim table = NumericTable.FromRows(tokens, rows, colnames)

call console.WriteLine($"word vectors: {table.nsamples} tokens x {table.nfeatures} dims")

' ---------------------------------------------------------------------------
' 2. Use UMAP to reduce the word vectors to 9 dimensions
'
'    The umap method accepts the unified 2D table and returns a new table whose
'    features are the embedding coordinates dim_1..dim_n; the row names and the
'    existing label columns are inherited
' ---------------------------------------------------------------------------
dim manifold = table.umap(dims := 9, neighbors := 64)

call console.WriteLine($"umap embedding: {manifold.nsamples} samples x {manifold.nfeatures} dims")

' ---------------------------------------------------------------------------
' 3. Run KMeans clustering on the reduced table
'
'    The clustering result is written into the label matrix of the table as the
'    label column cluster
' ---------------------------------------------------------------------------
dim clusters = manifold.kmeans(k := 9)

' ---------------------------------------------------------------------------
' 4. Extract the data needed for plotting from the result table
' ---------------------------------------------------------------------------
dim x = manifold.Feature("dim_1")
dim y = manifold.Feature("dim_2")
dim class_id = clusters.ClusterLabels()
dim classes(class_id.length - 1) as string
dim sizes(8) as integer

for i = 0 to class_id.length - 1
    classes(i) = class_id(i).ToString()
    sizes(class_id(i) - 1) += 1
next

call console.WriteLine($"kmeans cluster sizes: {String.Join(", ", sizes)}")

call SkiaDriver.Register()

Using plt As New ScatterPlot(800, 600, PlotTheme.Nature())
    plt.Title = "UMAP group of Rapunzel"
    plt.SubTitle = "UMAP scatter of the 'Rapunzel' word2vector embedding result"
    plt.XLabel = "UMAP1"
    plt.YLabel = "UMAP2"
    plt.Plot(DataSerials(x, y, classes).tolist())
    plt.SavePng(here("rapunzel-umap-groups.png"), 300)
End Using

' ---------------------------------------------------------------------------
' 5. Export the clustering result table as csv
'
'    The exported layout also follows the convention of
'    "row names + feature columns + label: prefixed label columns", so it can
'    be loaded back losslessly through NumericTableIO.ReadCsv
' ---------------------------------------------------------------------------
call clusters.WriteCsv(here("rapunzel-umap-groups.csv"))

call console.WriteLine("done: rapunzel-umap-groups.png")
call console.WriteLine("done: rapunzel-umap-groups.csv")

04 Results

The word-vector manifold

UMAP scatter of the Rapunzel word2vec embedding, colored by k-means cluster
Fig. 1 — First two UMAP dimensions of the 30-D CBOW embedding; every point is one of the 461 vocabulary words, colored by its k-means cluster. The manifold keeps distributionally similar words adjacent — the function-word cloud (and, had, long, to…) holds the left side, content-word neighborhoods weave through the middle, and recurring word families settle along the lower arc.
The k-means run is not a semantic claim — with k = 9 it simply partitions the embedding to make the manifold's shape visible. The interesting structure is the smooth word-neighborhood texture underneath, which is what the UMAP coordinates preserve.

05 Table preview

rapunzel-umap-groups.csv

461 vocabulary words × 9 UMAP dimensions + cluster label. The first and last rows of the exported table:

worddim_1dim_2dim_3dim_4dim_5dim_6dim_7dim_8dim_9label:cluster
There3.56365143724127-2.0734000598497624-1.6198605900202898-0.14376740383342634-2.4854460625022785-1.006297590081116-3.6068850016078593-1.6569097646708442-1.43511870486846154
were3.0914420761104258-1.958015080620244-2.06118647971116431.0601210819504878-2.421050814004005-0.5451589911057511-2.5416505018427222-2.127647654171174-0.71819749805856683
once3.289039017936846-1.4619525008248608-1.65663859453973330.11159778851742302-2.247780361752952-1.594633756897961-2.9916830355064667-1.2217044701082365-0.55461906976990691
a2.887581462689212-2.071979678993014-1.30204982930351190.543122471492502-2.120513215470194-1.2916062091648675-3.0632017744169233-0.7273641103968851-1.19375631984983989
···
afterwards2.7490901765149736-2.0951038677069533-1.76380457241878050.05883449852269965-1.802902414524226-1.204151837311557-2.3851538916094195-1.94099085443912-0.56679528360385096
happy2.3074653272538312-1.665422861944701-0.99343136177132610.49274067258745763-2.265759595867058-1.614801199608761-2.7155739136551-1.7584221848531272-1.6214090681636976
contented3.178361223400942-1.81659696214847-2.1061329581439043-0.33406266420931685-2.3633787915082607-0.5747379469595877-2.7257357726893594-1.0757773302040399-1.15688946935939944

Cluster sizes

ClusterWordsFirst members
Cluster 467There, man, wished, child, hoped, was’…
Cluster 355were, desire, garden, no, dreaded, this’…
Cluster 152once, and, woman, who, little, their’…
Cluster 960a, length, of, from, beautiful, flowers’…
Cluster 755had, vain, the, that, about, her’…
Cluster 836long, in, God, be, surrounded, high’…
Cluster 234for, to, These, one, an, enchantress’…
Cluster 660At, window, back, house, splendid, by’…
Cluster 542bed, Ah, his, salad, ate, him’…