This script analyses word frequency across isiXhosa text corpora (Leipzig, SADiLaR, Bible, NCHLT). It identifies the top 20 most frequent words per corpus, saves results to Excel, and visualises the results using Jaccard and cosine similarity heatmaps.