Writing

Notes on what I'm learning, in my own words.

121 posts found

/AI/4 min read

Attention Sinks

Attention weights must sum to exactly one, so "nothing here is relevant" cannot be expressed. The first token becomes where the surplus goes — and dropping it makes every remaining weight 2.63× sharper than the model intended.

  • attention-sinks
  • softmax
  • streaming

/AI/4 min read

Sliding Window Attention

Limiting each token to a 4,096-position window cuts the attention cost 32× at 131k tokens — and cuts the memory a sequence holds from 17.18 GB to a fixed 537 MB, which is 3 concurrent requests becoming 122.

  • attention
  • long-context
  • kv-cache

/AI/4 min read

TensorRT-LLM

A decode step launches 384 tiny kernels and spends 1.92 ms of CPU time just asking for them. Replaying the whole step as one recorded graph makes that 5 microseconds, which is 1.27× more tokens per second for no change in arithmetic.

  • tensorrt-llm
  • inference
  • cuda-graphs

/AI/4 min read

LPU

Putting the weights in on-chip SRAM makes a single stream 27× faster and costs 732× the throughput per chip. That one exchange rate explains every design decision in the architecture.

  • lpu
  • inference
  • sram

/AI/4 min read

Decoding EAGLE

Guessing the model's next hidden state instead of its next token removes the sampling noise from the guess. Going from 2.5 to 3.5 accepted tokens per pass is only an 11-point accept-rate gain, and a 40% speedup.

  • eagle
  • speculative-decoding
  • inference

/AI/4 min read

Decoding Medusa

Extra prediction heads on the same model replace a separate draft model — 134 MB instead of 2 GB. The acceptance chain multiplies, so the fourth head adds 0.044 tokens and the sixth adds 0.001.

  • medusa
  • speculative-decoding
  • inference

/AI/4 min read

RNNs vs Transformers

Attention costs n²d and recurrence costs nd², so they cross exactly at n = d — 768 tokens for a typical model. Transformers won on training parallelism, and pay for it past that crossover.

  • rnn
  • transformers
  • attention

/AI/4 min read

Definition of Done

A loop stops when its checker passes, so the checker sets a ceiling on quality that no number of attempts can lift. Tightening the checker buys more than improving the model does.

  • verification
  • agents
  • evaluation

/AI/4 min read

Prefix Tuning

Prefix tuning trains 48× more parameters than prompt tuning, and that factor is exactly twice the layer count. What the extra numbers buy is influence at every depth rather than only at the input.

  • prefix-tuning
  • peft
  • attention

/AI/5 min read

Graph Engineering

A graph with one three-way decision and a cycle capped at three has fifteen possible paths. The same work as a free loop with six tools and twelve steps has 2.61 billion. Only one of those numbers can be tested.

  • graphs
  • agents
  • testing

/AI/5 min read

Loop Engineering

An agent's cost grows with the square of its step budget, so doubling the budget nearly quadruples the bill. In one accounting the last six of twelve turns served 12% of runs and carried 71% of the tokens.

  • agents
  • loops
  • context

/AI/4 min read

Lost in the Middle

Accuracy follows a U-shape across a long input, which means the order you pass chunks in is worth real points. Ten chunks arranged well beat twenty in rank order, using half the tokens.

  • long-context
  • rag
  • retrieval

/AI/5 min read

LLM Watermarking

Detecting a watermark is a coin-flip test. At full strength 64 tokens is enough; every halving of the signal quadruples the text you need, which is exactly why paraphrasing defeats it.

  • watermarking
  • detection
  • statistics

/AI/5 min read

Deep RL from Human Preferences

Nine hundred comparisons — one hour of someone's attention — taught a simulated robot to backflip. That is 11,111× fewer human decisions than scoring every step, and the trick is what the comparison is asked about.

  • rlhf
  • reward-modelling
  • reinforcement-learning

/AI/5 min read

Decoding InstructGPT

Thirteen thousand written examples is 0.002% of the pretraining text, and the whole alignment run was 1.6% of the compute. It made a 1.3-billion-parameter model preferred over a 175-billion one.

  • instructgpt
  • rlhf
  • alignment

/AI/4 min read

Decoding ColBERT

Keeping one vector per token instead of one per passage lets every query term find its own best match. On a worked example that widened the gap between the right document and a plausible wrong one from 64% to 113%.

  • colbert
  • late-interaction
  • maxsim

/AI/4 min read

LLM Guardrails

A guardrail catching 90% of harmful requests with a 1% false-positive rate blocks 1,267 requests a day, of which 997 are legitimate. The visible effect of most guardrails is refusing real users.

  • guardrails
  • safety
  • false-positives

/AI/5 min read

Cloud vs On-device

In the cloud the data travels to the model; on a device the model travels to the data. Three numbers decide which: the latency budget, the crossover request count, and the fraction of users who never take your update.

  • deployment
  • on-device
  • latency

/AI/4 min read

Encoder vs Decoder

One mask setting separates the two, and it decides two things that look unrelated: how much training signal a corpus yields, and whether generation can be made incremental. Both favour the decoder by a wide margin.

  • transformers
  • encoder
  • decoder

/AI/4 min read

Generative AI

A classifier picks one of a handful of labels. A generator picks one of 10^470 possible answers, which is why it cannot work by choosing from a list — and why its characteristic failure is being fluent and wrong.

  • generative-ai
  • llm
  • autoregression

/AI/5 min read

Prompt Injection in LLMs

A filter that blocks 95% of attempts sounds strong and falls to a coin flip in fourteen tries. Which is why the defences that work are the ones enforced in code, where the number is 100% by construction.

  • prompt-injection
  • security
  • agents

/AI/4 min read

Precision vs Recall

The trade-off everyone describes is real and it is not the main problem. When positives are rare, a 98%-accurate classifier produces alerts that are wrong 91% of the time — and the fix is not a better model.

  • precision
  • recall
  • base-rate

/AI/4 min read

Embeddings

Two numbers can hold two distinguishable concepts. Sixteen can hold twenty-one. That measured curve is the reason real embeddings have hundreds of dimensions, and it is rarely the explanation given.

  • embeddings
  • vectors
  • cosine-similarity

/AI/4 min read

Agent Skills

Loading fifty procedures eagerly costs 210,000 tokens and does not fit. Loading their descriptions costs 4,750, which makes the whole thing work — and turns the real problem into a fifty-way choice made from twenty words.

  • agent-skills
  • context
  • progressive-disclosure

/AI/4 min read

MCP

A protocol turns an N×M integration problem into N+M, which is 8.6× less work at a dozen apps and thirty systems. The part that needs more attention is that a server's own text reaches the model's context.

  • mcp
  • protocol
  • agents

/AI/4 min read

Fine-tuning

Fine-tuning is good at changing how a model behaves and bad at teaching it facts. It also has a break-even point — about 6,750 requests in one worked case, and ten times that if the prompt it replaces was being cached.

  • fine-tuning
  • lora
  • training

/AI/4 min read

OKF

An agent can work out a column's type by looking. It cannot work out that status 'H' means excluded from revenue, because that fact is in someone's head — and four rules out of five are like that.

  • okf
  • metadata
  • agents

/AI/4 min read

Semantic Search

Matching by meaning solves the vocabulary problem and creates a new one: one vector per chunk averages everything together, so a single distinguishing term contributes about 1/√n of the direction. At 200 tokens that is 7%.

  • semantic-search
  • embeddings
  • cosine-similarity

/AI/5 min read

Cursor

Three different models do three different jobs, and the reason is a latency budget. A completion has about 210 ms of model time before it feels late, and a large model needs six times that.

  • cursor
  • code-editors
  • latency

/AI/4 min read

Context Compaction

Replacing old turns with a summary buys room, and each summary is lossy. Summarise a summary eight times and 43% of the original detail is left — which is why the facts that matter must never be summarised twice.

  • context-compaction
  • summarisation
  • agents

/AI/4 min read

Claude Code

A coding agent is a model plus tools plus a loop. The two things that make it work are searching instead of reading — a 700× reduction in what enters the context — and being able to check its own answer, which turns 55% into 96%.

  • claude-code
  • coding-agents
  • tools

/AI/4 min read

llama.cpp

Running a model on a laptop CPU is a memory-bandwidth problem, not an arithmetic one — which sets a hard ceiling of about 12 tokens a second. It also explains why offloading half the layers to a GPU barely helps.

  • llama-cpp
  • local-inference
  • cpu

/AI/4 min read

Model Quantization

Fewer bits per number is the easy part. What decides whether it works is how many numbers share a scale — and in one measured case the average error looked fine at 5% while the ordinary features were off by 42%.

  • quantization
  • int8
  • outliers

/AI/4 min read

Chain-of-Thought Prompting

A model has no scratchpad except the tokens it has already written, and each token gets one fixed pass through the network. Asking for the steps is asking for more computation — sixty tokens of working is sixty times as much.

  • chain-of-thought
  • prompting
  • reasoning

/AI/4 min read

Prompt Chaining

Splitting a task across several prompts makes each step more reliable and makes more steps that all have to work. Those pull against each other, and the result is an optimal chain length that is shorter than most people assume.

  • prompt-chaining
  • prompting
  • reliability

/AI/4 min read

PyTorch

PyTorch builds the computation graph by running your code, which is why ordinary Python control flow works inside a model. The same decision explains its two most common bugs.

  • pytorch
  • autograd
  • tensors

/AI/4 min read

Semantic Caching

An exact cache is never wrong. A semantic cache answers a question nobody asked, which makes the similarity threshold a decision about how often you are willing to be confidently incorrect.

  • semantic-caching
  • embeddings
  • similarity

/AI/4 min read

Hybrid Search

Keyword and vector search fail on opposite queries, so run both. The hard part is combining two ranked lists whose scores are not on the same scale — and the constant in the standard fix quietly decides your ordering.

  • hybrid-search
  • bm25
  • rrf

/AI/4 min read

HyDE in RAG

A question and its answer barely resemble each other, which is a problem when retrieval works by resemblance. HyDE has the model write a fake answer first and searches with that instead.

  • hyde
  • rag
  • retrieval

/AI/4 min read

Prefill vs Decode

Serving a model is two workloads wearing one name, and their metrics pull against each other. Time to first token is 1.5% of a long answer's duration and almost all of how responsive it feels.

  • prefill
  • decode
  • latency

/AI/5 min read

GPU for Deep Learning

A GPU is thousands of simple cores, which suits the arithmetic of a neural network perfectly. What actually decides what you can build on one is not the cores — it is that training a model needs eight times the memory of running it.

  • gpu
  • vram
  • training

/AI/4 min read

LangGraph

The point is not that graphs are more flexible than lines. It is that the graph's position is stored alongside its data, so a run can be paused, resumed after a crash, or held for a human — none of which a straight-line pipeline can do.

  • langgraph
  • agents
  • state-machines

/AI/4 min read

LangChain

The framework's actual content is one interface that every piece implements. That is what makes the pipe operator work, what makes batching and streaming propagate through a chain you assembled — and what makes a broken prompt hard to find.

  • langchain
  • llm
  • composition

/AI/4 min read

SGLang

SGLang keeps every in-flight request's attention cache in one prefix tree, so any overlap between any two requests is reused automatically — including overlaps nobody declared. Its second trick is not calling the model for output that was never in doubt.

  • sglang
  • radix-tree
  • serving

/AI/4 min read

Approximate Nearest Neighbour Search

Giving up the guarantee of finding the true nearest vector buys enormous speed for almost no accuracy. In a measured run, scanning 3.4% of the data found the exact answer 99.8% of the time.

  • ann
  • vector-search
  • hnsw

/AI/4 min read

Google TPU

A TPU is built around a grid of multipliers that pass values to each other instead of fetching them. That one decision is where its speed and its power efficiency come from — and also where its awkwardness comes from.

  • tpu
  • systolic-array
  • hardware

/AI/3 min read

Image Embeddings

Comparing images by their pixels fails badly — the same object moved twelve pixels can look less like itself than a different object does. An embedding is a representation trained so that stops happening.

  • image-embeddings
  • computer-vision
  • similarity

/AI/4 min read

Computer-Use Agents

An agent that drives a screen with a mouse and keyboard needs no integration at all, which is its whole appeal. It also has to be right at every single step, and that requirement is much harsher than it sounds.

  • computer-use
  • agents
  • vision

/AI/4 min read

Decoding Sakana Fugu

Fugu does not answer questions. It decides which frontier model should, and in its larger form how several of them should divide the work — which makes the interesting question how much capability is actually available to be routed to.

  • orchestration
  • routing
  • multi-model

/AI/3 min read

Diffusion Language Models

Instead of writing left to right, a diffusion language model starts from an all-masked sequence and fills it in over a few rounds. The speedup is real, and so is the reason it cannot be pushed as far as it looks.

  • diffusion
  • language-models
  • parallel-decoding

/AI/4 min read

Embedding Cache

The same text embeds to the same vector every time, so recomputing it is pure waste. What makes the cache worth building is that query traffic is wildly lopsided — a few megabytes of memory removes most of the calls.

  • embeddings
  • caching
  • rag

/AI/4 min read

World Models

A world model is a learned simulator an agent can practise inside. The whole design question is how many steps it can imagine before its own small errors compound into fiction.

  • world-models
  • reinforcement-learning
  • planning

/AI/4 min read

vLLM

vLLM's central idea is to stop giving each request a contiguous slab of GPU memory and instead hand it small blocks through a lookup table — which removes the waste, and as a side effect makes sharing memory between requests almost free.

  • vllm
  • paged-attention
  • serving

/AI/5 min read

LLM Inference Optimization

Generating a token barely uses the GPU's arithmetic — it spends almost all its time moving weights from memory. Once you see that, every serving optimization sorts itself into one of three piles.

  • inference
  • gpu
  • memory-bandwidth

/AI/4 min read

Function Calling in LLMs

The model never runs anything. It writes a structured request and stops, and your code decides whether to honour it — which is both why function calling is safe and where every mistake with it comes from.

  • function-calling
  • tools
  • json-schema

/AI/4 min read

GGUF

GGUF is one file holding the weights, the tokenizer and everything needed to run a model — laid out so the runtime can map it into memory instead of parsing it, and quantized in small blocks so 4 bits per weight is actually usable.

  • gguf
  • quantization
  • local-inference

/AI/4 min read

Knowledge Distillation

A trained model knows more than the answer it gives — it knows how close the wrong answers were. Distillation trains a small model on that whole distribution instead of the single correct label, which is a far richer target.

  • knowledge-distillation
  • training
  • softmax

/AI/4 min read

Token Streaming

A model produces its answer one token at a time, so the answer exists in pieces long before it is finished. Streaming is the decision to send those pieces instead of holding them, and it works because reading is far slower than generating.

  • token-streaming
  • sse
  • llm-inference

/AI/4 min read

Prompt Caching

Most of a long prompt is the same on every request. Prompt caching stores the work already done on that unchanging opening so the model never has to redo it — but only if the opening really is unchanged, character for character.

  • prompt-caching
  • kv-cache
  • llm-inference

/AI/4 min read

JEPA

Predict what a hidden part of an image means rather than what it looks like. The objective that makes that work has a perfect and useless solution, and the architecture exists to keep it out of reach.

  • jepa
  • self-supervised
  • representation

/AI/4 min read

Rerankers

Retrieval is fast because it never reads the question and the document together. A reranker does read them together, which is why it is better and why it can only be used on a shortlist.

  • reranking
  • retrieval
  • rag

/AI/5 min read

Vector databases

Storing meaning as coordinates is the easy half. The hard half is finding the nearest ones among a hundred million without reading all of them — which is why the answer is allowed to be slightly wrong.

  • vector-database
  • ann
  • embeddings

/AI/4 min read

Dropout

Switch off a random half of the neurons at every training step. It stops the network relying on any particular one, and the scaling factor that makes it work is more interesting than it looks.

  • dropout
  • regularisation
  • overfitting

/AI/4 min read

GANs

Two networks trained against each other, one making fakes and one catching them. Because each is graded by the other, the loss numbers stop meaning what loss numbers usually mean.

  • gan
  • generative
  • adversarial

/AI/5 min read

Continual learning in LLMs

Teaching a trained model something new tends to erase what it already knew. Every fix for that is a different way of deciding which weights are allowed to move.

  • continual-learning
  • forgetting
  • fine-tuning

/AI/5 min read

Variational autoencoders

An ordinary autoencoder compresses well and generates nothing, because its latent space is full of holes. A VAE fills the holes by encoding to regions instead of points — and pays for it with a second loss term.

  • vae
  • latent-space
  • generative

/AI/4 min read

Diffusion models

Destroy an image with noise in a thousand small steps, then train a network to undo one step at a time. Each individual step is easy, which is the entire reason the method works.

  • diffusion
  • generative
  • noise-schedule

/AI/4 min read

Multi-head attention

Running several attention operations side by side costs nothing extra — the parameters and the cache are identical to one big head. What you actually spend is the width each head gets to work in.

  • attention
  • transformers
  • heads

/AI/4 min read

Cross attention

One sequence asks the questions and a different sequence answers them. Two useful things follow from that split: the attention matrix stops being square, and the answering side is computed once.

  • attention
  • transformers
  • encoder-decoder

/AI/4 min read

Self attention

The word self means the queries, keys and values all come from the same sequence. The consequence is that a word's representation is built from its neighbours, so the same word comes out different in different sentences.

  • attention
  • transformers
  • context

/AI/5 min read

AI agent observability

An agent's answer is the smallest part of what it did. Recording the rest produces far more data than ordinary logging — around seventy times more — which makes what you keep a real decision.

  • observability
  • agents
  • tracing

/AI/4 min read

How AI agents communicate

Once several agents work on one task, they have to pass things to each other. The shape of who-talks-to-whom is the decision that matters, and the arithmetic makes it for you.

  • agents
  • messaging
  • protocols

/AI/4 min read

AI subagents

A subagent is an agent called like a function: it runs its own loop, returns a value, and throws its working away. That discarded working is the entire point.

  • agents
  • subagents
  • context

/AI/4 min read

The agent loop

The loop is about twenty lines of ordinary code. Its cost is not ordinary — because the whole history is re-sent every turn, tokens grow with the square of the turn count.

  • agents
  • loop
  • cost

/AI/4 min read

AI orchestration

Wiring several model calls and tools into one workflow. The shape you choose is not a style preference — it decides latency, cost, and how much you can debug.

  • orchestration
  • workflows
  • latency

/AI/4 min read

AI agent evaluation

Judging an agent by whether it finished is not enough, because per-step reliability compounds. A 95%-per-step agent completes a twenty-step task about a third of the time.

  • agents
  • evaluation
  • trajectory

/AI/4 min read

LLM evaluation

There is no formula for whether an answer was any good, so evaluation becomes a choice about what to trade away. The costs differ by two orders of magnitude, which is what decides the shape of a real setup.

  • evaluation
  • benchmarks
  • metrics

/AI/4 min read

LLM as a judge

Using one model to grade another scales in a way human review cannot. It also brings biases that are measurable — and at least one of them cancels exactly if you run the comparison twice.

  • evaluation
  • judge
  • bias

/AI/4 min read

Contrastive learning

Teach a model what things mean by showing it what is the same and what is different. Almost all the learning signal comes from the few negatives it nearly got wrong.

  • contrastive
  • embeddings
  • self-supervised

/AI/4 min read

Recursive language models

Rather than reading a huge input, the model writes code that reads it — calling other models on the pieces and seeing only what comes back. What that buys is not fewer tokens.

  • long-context
  • recursion
  • agents

/AI/5 min read

GRPO

Generate several answers to the same question, and score each one against how the others did. The group's own average replaces a whole second network — and it fails in a way that is easy to measure.

  • grpo
  • reinforcement-learning
  • baseline

/AI/4 min read

DPO

DPO removes the reward model by noticing that a language model already contains one. What it optimises is a difference, and that turns out to matter more than it sounds.

  • dpo
  • alignment
  • preferences

/AI/4 min read

PPO

PPO improves a policy while refusing to move it far in one step. The clipping that enforces this is not symmetric, and the asymmetry is the part worth understanding.

  • ppo
  • reinforcement-learning
  • clipping

/AI/4 min read

BatchNorm vs LayerNorm

Same formula, different axis. One averages a feature across the examples in a batch, the other averages all the features within one example — and that choice decides which architectures each one suits.

  • normalisation
  • batchnorm
  • layernorm

/AI/5 min read

Large reasoning models

A reasoning model spends thousands of tokens working before it answers. Those tokens are billed, they set the latency, and they are the reason it gets harder problems right.

  • reasoning
  • test-time-compute
  • rlvr

/AI/5 min read

RLHF

People cannot score an answer out of ten consistently, but they can say which of two is better. RLHF is the machinery for turning a pile of those comparisons into a model that behaves.

  • rlhf
  • alignment
  • reward-model

/AI/4 min read

Autoregressive models

Generate one token, append it to what you have, and use that to generate the next. The whole design follows from one identity in probability, and so does its main weakness.

  • autoregressive
  • generation
  • chain-rule

/AI/4 min read

Continuous batching

Waiting for every request in a batch to finish leaves most of the GPU's slots idle most of the time. Refilling each slot the moment it empties is a scheduling change, and it roughly doubles what the same hardware serves.

  • inference
  • batching
  • throughput

/AI/5 min read

Small language models

Under about ten billion parameters, a model stops needing a datacentre. What actually decides where it can run is not the parameter count but how many gigabytes it occupies.

  • slm
  • quantisation
  • on-device

/AI/4 min read

Multimodal AI

A picture and the sentence describing it are completely different kinds of data. Multimodal models work by turning both into vectors that land in the same place.

  • multimodal
  • embeddings
  • vision

/AI/5 min read

LLM routing

Sending every request to the largest model means paying premium prices for questions a small model answers perfectly. Routing picks a model per request — and it tolerates being wrong far more than you would expect.

  • routing
  • cost
  • inference

/AI/4 min read

Context engineering

Everything a model knows about your task is in one window, and the pieces compete for space. Deciding what goes in, in what order, and what gets dropped is most of the work.

  • context
  • prompting
  • rag

/AI/4 min read

Reflection agents

The agent writes something, criticises its own draft, and rewrites. It works because judging is an easier job than writing — and it fails in a specific way when the critic asks for facts it cannot check.

  • agents
  • reflection
  • critic

/AI/5 min read

Speculative decoding

A small model guesses the next few tokens and a large one checks them all in a single pass. Done properly the output is not an approximation of the large model — it is exactly what the large model would have produced.

  • inference
  • speculative-decoding
  • latency

/AI/4 min read

GraphRAG

Chunk search finds passages that look like the question. When the answer is a chain running through documents that mention none of the question's words, that is not enough.

  • rag
  • knowledge-graph
  • retrieval

/AI/4 min read

Plan-and-execute agents

Instead of deciding one step at a time, the agent writes the whole plan first and then works through it. Cheaper and steadier when the path is knowable, and useless when it is not.

  • agents
  • planning
  • execution

/AI/4 min read

Agentic RAG

Ordinary retrieval searches once and hopes the result was good. Agentic RAG puts an agent in charge of searching, so it can judge what came back, rewrite the query, and go again.

  • rag
  • retrieval
  • agents

/AI/4 min read

ReAct agents

ReAct makes an agent write down its reasoning before every action. The thought is not decoration — it is what the model reads back on the next step, and it is the only place you can see why anything happened.

  • agents
  • react
  • reasoning

/AI/5 min read

Multi-agent systems

Splitting a job across several specialised agents buys parallelism and focus, and costs latency, tokens and the ability to debug easily. It is worth it less often than it looks.

  • agents
  • multi-agent
  • orchestration

/AI/5 min read

Decoding DeepSeek-V4

A million tokens of context is unaffordable with ordinary attention. DeepSeek-V4 gets there by summarising the distant past into far fewer entries and attending to those instead.

  • deepseek
  • moe
  • attention

/AI/4 min read

AI agent memory

A language model remembers nothing between calls. Everything that looks like memory is machinery around it deciding what to write down and what to put back in front of the model.

  • agents
  • memory
  • context-window

/AI/4 min read

AI agents

An agent is a language model wrapped in a loop that lets it call tools and see what happened. The loop is the whole difference between answering a question and doing a job.

  • agents
  • tools
  • llm

/AI/4 min read

RMSNorm

Layer normalisation does two things: it recentres values and it rescales them. RMSNorm drops the recentring, keeps the rescaling, and modern language models have almost all switched to it.

  • rmsnorm
  • layernorm
  • normalisation

/AI/4 min read

LoRA

Fine-tuning normally means updating every weight in the model. LoRA freezes them all and learns a pair of thin matrices alongside, which turns out to be enough.

  • lora
  • fine-tuning
  • adapters

/AI/5 min read

The math behind RoPE

RoPE gives a transformer position information by rotating query and key vectors instead of adding anything to them. The rotation cancels in the dot product, leaving only the gap between two tokens.

  • rope
  • position-embedding
  • attention

/AI/4 min read

Grouped query attention

Every attention head keeping its own keys and values makes the cache enormous. Grouped query attention has heads share them in small groups, cutting the memory by the size of the group.

  • attention
  • gqa
  • kv-cache

/AI/5 min read

The math behind cross-entropy loss

Cross-entropy turns a set of predicted probabilities into one number saying how wrong they were. It is built so that being confidently wrong costs far more than being unsure.

  • cross-entropy
  • loss
  • softmax

/AI/5 min read

The math behind gradient descent

Training a model means finding the weights that make the error smallest. Gradient descent does it by repeatedly stepping downhill, and the size of those steps decides whether it works at all.

  • gradient-descent
  • training
  • learning-rate

/AI/4 min read

Vision transformers

A vision transformer runs an image through the same machinery a language model runs text through. Almost all of the work is in turning a picture into something shaped like a sentence.

  • vit
  • transformers
  • computer-vision

/AI/4 min read

Feed-forward networks in LLMs

Attention gets the attention, but two thirds of a transformer layer's parameters sit in the feed-forward network. It is where the model keeps what it knows.

  • llm
  • transformers
  • feed-forward

/AI/5 min read

Decoding flash attention

Flash attention produces exactly the same numbers as ordinary attention. It is faster because it never writes the score matrix to memory at all.

  • attention
  • gpu
  • flash-attention

/AI/4 min read

Mixture of experts in LLMs

A mixture-of-experts layer holds many small networks and sends each token through only two of them. The model gets to be large without every token paying for all of it.

  • llm
  • moe
  • router

/AI/5 min read

Decoding the transformer architecture

A transformer takes tokens in and gives tokens out. Between those two ends sits a short list of components, each with one job — and the same list underpins BERT, GPT and everything since.

  • transformers
  • architecture
  • encoder

/AI/5 min read

The math behind backpropagation

Training a network means working out how much each weight contributed to the error. Backpropagation does it with one rule from calculus, applied backwards through the network.

  • neural-networks
  • backpropagation
  • chain-rule

/AI/3 min read

Why attention divides by √dₖ

The scaling factor in scaled dot-product attention is not a tuning knob. It is the one number that keeps the attention scores at a usable size no matter how wide the vectors are.

  • llm
  • attention
  • softmax

/AI/4 min read

The math behind attention: Q, K and V

Attention is one formula built from three matrices. Working it end to end on a three-word sentence shows exactly what Q, K and V are doing.

  • llm
  • attention
  • transformers

/AI/4 min read

Harness engineering in AI

A model on its own can only turn text into more text. The harness is everything you build around it — tools, memory, guardrails, retries — and it is usually where most of the work lives.

  • llm
  • agents
  • harness

/AI/5 min read

Byte pair encoding in LLMs

A model cannot read text, only numbers, so text has to be cut into pieces first. Byte pair encoding decides where to cut by repeatedly merging the most common pair of neighbours.

  • llm
  • tokenization
  • bpe

/AI/5 min read

Paged attention in LLMs

Serving an LLM means reserving memory for a reply before you know how long it will be. Paged attention stops that reservation from wasting most of the GPU.

  • llm
  • inference
  • paged-attention

/AI/3 min read

KV cache in LLMs

An LLM generates one token at a time, and each new token needs the keys and values of every token before it. The KV cache stores them instead of recomputing them.

  • llm
  • inference
  • kv-cache

/AI/4 min read

Causal masking in attention

Causal masking is what stops a token from attending to the tokens that come after it. Without it, a language model can see the answer it is being asked to predict.

  • llm
  • attention
  • transformers

Writing next

Android

  • Compose performance: when to use derivedStateOf vs remember
  • Modularising a multi-million-line Android monorepo
  • On-device LLM inference with Foundation Models

AI

  • Building an autonomous job-application agent with LangGraph
  • Prompt caching cost math: when it actually pays off
  • Playwright as a tool surface for agents
Back home