-
From Words to Attention: The Full Story of How AI Learned to Represent Language (Summary)
A single throughline connects one-hot encoding, word2vec, fastText, and the attention mechanism inside every modern LLM — each one solving the exact problem the last one left behind. Here’s the…
-

The “Bank” Problem: Why Your RAG System Might Be Retrieving the Wrong Thing
A word embedding for “bank” has to mean both a riverbank and a financial institution at once. Here’s why that quiet compromise causes real retrieval errors in production RAG and…
-
Subword Embeddings and fastText: Solving the OOV Problem at the Vector Level (Ep.01.01)
Episode 01.00 ended by asking: does Module 00’s word-embedding machinery just apply directly to subword tokens like lowe + r + i + n + g, or does something break?…
-
Why Are “Unrelated” Word Embeddings Never Actually at Zero Similarity? (Ep. 00.04)
Episode 00.03 ended with similarity(“king”, “apple”) = 0.74 — high, even though apple is the clearly unrelated word in that toy vocabulary — and asked why. There are two real…
-
Dense Word Embeddings Explained: From Co-occurrence Matrices to word2vec (Episode 00.03)
The last episode 00.02 left a real problem on the table: co-occurrence vectors give words meaningful, graded similarity — but they’re the same size as the vocabulary, mostly zeros, and…
-
How Does a Word Become a Number? The Representation Problem in AI (Episode 00.02)
Episode 00.01 landed on a working definition of intelligence built around selecting good actions under uncertainty, and generalizing efficiently. Every example we used to demonstrate that — the lookup table,…
Search
Categories
- Artificial Intelligence (AI) (46)
- Learn FastAPI from Scratch (12)
- Shopify (1)
- Uncategorized (97)
- Versana Companion Plugin (7)
- Weebly (24)
- Weebly Apps (7)
- Wordpress (58)
- WordPress Block Theme Development (20)
- Zero to Automation Pro: The Complete n8n Developer Series (3)
Recent Posts
- How LSTMs Fix the Vanishing Gradient Problem: The Gating Math, Derived and Measured (Ep:04.01)
- RNNs and the Vanishing Gradient Problem: Why Attention Had to Be Invented (Ep:04.00)
- Why Your Classifier Trains So Slowly at First (and MSE Might Be Why)
- How a Neural Network Actually Learns: From a Single Neuron to Backpropagation
- Why Cross-Entropy Beats MSE for Classification: The Gradient Math, Proven (Ep:03.06)
Tags
AI course AI engineering angular backpropagation block theme BlockTheme c programming cross-entropy loss data science deep learning fastapi FastAPI development environment full site editing gradient descent gutenberg java learn AI from scratch learn html & html5 learn wordpress machine learning math for machine learning n8n beginner guide n8n tutorial neural networks nginx NLP fundamentals nodejs react R programming vanishing gradient versana versana companion plugin vps server website migration weebly app weebly apps weebly tutorials word embeddings wordpress blocks wordpress block theme wordpress news wordpress plugin development wordpress plugins wordpress security wordpress tutorials