Week 20. Multimodality and long context
Learn in the app: tutor, coding problems →
Core: ViT, CLIP and the contrastive loss, the LLaVA projector, RoPE extrapolation · Depth: cross-attention and the Q-Former, diffusion (05-ГЛУБИНА) · ≈ 10 h core / 22 h total
Two topics with a common root: a transformer does not care where the vectors of the sequence came from. Multimodality answers the question **how do you turn an image or a sound into tokens. Long context answers the question how do positions behave where the model has never seen them**.
Part 1. ViT
In plain terms. A 224×224 image is cut into 16×16 squares: 14 across, 14 down, 196 patches in total.
Each patch holds 16·16·3 = 768 numbers, and a linear layer turns them into a vector of size D.
From there the image becomes a sentence of 196 "words", and the transformer processes it like text.
A 448×448 image already gives 784 "words", four times as many, and 784² / 196² = 16 times as many pairs for attention.
- The image is cut into
P×Ppatches, and each one is linearly projected toD. This is exactly a convolution with kernelPand strideP. At 224×224 andP = 16you get 196 tokens - The cost takeaway: resolution ×2 → tokens ×4 → attention ×16. That is why multimodal models fight so hard over the number of visual tokens
- Learned positional embeddings tie the model to a resolution (changing the resolution requires interpolation). 2D-RoPE (position = row and column) removes this tie
- Optional: write a NumPy convolution
conv2d(x, w, stride, padding)and check that with kernel and stridePand no padding it matches cutting into patches plus a linear layer. The output size⌊(n + 2·pad − P)/stride⌋ + 1, wherenis the image side, is asked about separately in interviews
Part 2. CLIP and contrastive learning
In plain terms. A batch of 4 pairs: 4 images and 4 captions. Compute the similarity of every image with every caption to get a 4×4 table. For image 1 the correct caption is the first one, and the other three serve as wrong options. Each row of the table is a 4-way classification, with the correct answer on the diagonal. Same for the columns: for each caption, find its image. With a batch of 32 768, as in CLIP, each row has 32 767 negatives.
- Contrastive learning: pull the representations of correct pairs together and push incorrect ones apart.
N(image, text) pairs → anN×Nsimilarity matrix, symmetric cross-entropy over rows and columns, the correct answer on the diagonal - Takeaway: this is classification where the classes are the other elements of the batch. The batch size equals the number of negatives, which is why CLIP is trained with huge batches. SigLIP replaces softmax with a sigmoid per pair and removes the normalization over the whole batch
Part 3. Connecting modalities
- Projector (LLaVA): an MLP maps the output of the vision encoder into the LLM's embedding space, and visual tokens are inserted into the sequence like ordinary ones
- Cross-attention (Flamingo): separate layers where the text attends to the image
- Q-Former (BLIP-2, "querying transformer"): a fixed number of learned queries compresses the image
into
Ktokens. The general trade-off: number of visual tokens versus detail - Audio: Whisper is an encoder-decoder over a log-mel spectrogram (sound energy across frequencies over time, on a logarithmic scale). Neural codecs (networks that compress sound into a stream of discrete codes) produce discrete audio tokens, so sound can be generated like text
Part 3b. Generation without autoregression
Images in multimodal systems are almost never drawn by autoregression. To see why, first let us say precisely what an LLM is.
In plain terms. Vocabulary {a, b}, length 2. The model outputs p(a) = 0.6 and p(b | a) = 0.3. Then p(ab) = 0.6·0.3 = 0.18.
The probability of any of the four strings is computed exactly, by multiplication.
- An LLM models an explicit factorization
p(x) = Π_{t=1..T} p(x_t | x_<t)(the chain rule, which holds for any distribution). The cross-entropy over positions (week 4) is exactly−log p(x). The cost: generation takesTsequential steps - The other families answer two questions differently: what the model learns, and how you get a sample from it
| Family | What it models | How it samples | Likelihood |
|---|---|---|---|
| Autoregression | probability of the next token given the prefix | token by token, T steps | exact |
| VAE (variational autoencoder) | an encoder to a latent z and a decoder from z ~ N(0, I) to data | take z, one decoder pass | lower bound (ELBO) |
| Diffusion | a denoiser: what noise was mixed in at level t | from pure noise, M denoising steps | lower bound |
| Flow matching | a velocity field v(x, t) from noise to data | solve an ODE, from 1 to tens of steps | via the ODE, expensive |
| GAN (generative adversarial network) | a generator with no density, plus a discriminator | one generator pass | none |
| EBM (energy-based model) | an energy U(x), p(x) ∝ exp(−U(x)) | MCMC (a random walk that converges to p), many steps | up to a constant |
The softmax from week 4 is an EBM on a finite set: the logits are negative energies, and the partition function can be computed.
Diffusion (DDPM, Ho et al., 2020). In plain terms. One "pixel" x₀ = 3, a noise level with ᾱ = 0.64 (the signal's share of the variance).
The noisy version: x_t = √0.64·3 + √0.36·ε = 2.4 + 0.6·ε, ε ~ N(0, 1). You draw ε = −1 and get x_t = 1.8.
The model sees 1.8 and the level index and guesses ε. It guesses −1: the clean value is recovered, (1.8 + 0.6)/0.8 = 3.
0.64 + 0.36 = 1, so for data with variance 1 the noisy version also has variance 1, and at the last level pure N(0, 1) remains.
- Forward process:
x_t = √ᾱ_t·x₀ + √(1 − ᾱ_t)·ε, whereᾱ_tdecreases from almost 1 to almost 0. Any level is reached in one step - Training: a random level, a random
ε, the loss‖ε − ε_θ(x_t, t)‖². Ordinary MSE regression, with no adversary and no MCMC - Generation: start from
N(0, I), thenMtimes (M: the number of steps) subtract the predicted noise and mix in a little fresh noise (DDPM uses hundreds of steps, fast samplers bring it down to tens). All positions are denoised at once; there is no "left to right" order - The condition (text) enters the denoiser through cross-attention. Classifier-free guidance: mix the predictions with and without the condition, and move further in the direction of the conditional one
Latent diffusion (Rombach et al., 2022). A 512×512×3 image is 786 432 numbers. An autoencoder (the VAE from the table) compresses it
into a 64×64×4 latent, that is, 16 384 numbers, 48 times fewer, and diffusion runs in the latents; the decoder returns pixels once.
The autoencoder takes care of fine details, and the expensive iterative part works with the essence. A related idea: VQ codes
(indices of vectors from a learned codebook) as "visual tokens" that can be generated autoregressively.
Masked diffusion for text (MDLM, Sahoo et al., 2024; LLaDA, Nie et al., 2025). In plain terms. For text the noise
is not Gaussian but a mask. At level t = 0.5, 4 of 8 tokens are masked, the model predicts all 4 at once, and CE is computed over the masked ones.
Generation: 8 masks, and over 4 steps you reveal 2 tokens at a time, the ones the model is most confident about.
- The difference from BERT (an encoder trained to guess masked tokens): there a fixed fraction is masked, about 15%, while here the level is random from 0 to 1, and the model can start from a fully masked string, which means it can generate. Pros: bidirectional attention, several tokens per step, editing the middle. Cons: no usual KV cache, and only a lower bound on the likelihood
Flow matching (Lipman et al., 2023; rectified flow, Liu et al., 2023). In plain terms. Noise z = −2, data x = 4.
Connect them with a straight line: x_t = (1 − t)·z + t·x; at t = 0.25 the point is −0.5, and the velocity along the whole path is x − z = 6.
The model v_θ(x_t, t) learns to predict this velocity with MSE. Generation: take noise and solve dx/dt = v_θ(x, t) with Euler's method.
For a single pair one step −2 + 6 = 4 would be enough, but the lines for different pairs cross, the model learns their average,
and the paths bend, so you need several steps. Diffusion can also be written as an ODE: flow matching just picks simpler paths.
- Takeaway: text holds on to autoregression because of exact likelihood, the KV cache and discreteness. Diffusion is strong where data is continuous and all positions are coupled (pixels, sound). Masked diffusion tries to win text back with parallel generation
Part 4. Long context
In plain terms. Head H = 128, base Θ = 10 000. The fastest RoPE pair rotates 1 radian per position:
a full circle every ~6 positions, and during training it sees every angle. The slowest one makes a full turn in about 54 000 positions.
When training on 4096 tokens it only manages to rotate ~27°. At position 16 384 the angle will be ~108°, which the model has never seen.
Position interpolation divides positions by 4: the angles are again within 27°,
but neighboring tokens now differ by a quarter step, including for the fast pairs.
- From week 7: the slow RoPE pairs do not complete a full turn during all of training. Takeaway: beyond the training length it is exactly these pairs that break: their angles are out of distribution, while the fast pairs have seen every angle
- Position interpolation: compress positions
m → m·L/L', so the angles stay within what was seen, but the fast frequencies lose resolution between neighbors. IncreasingΘ(NTK-aware) stretches the slow frequencies more than the fast ones. YaRN sets different rules for different frequency bands, plus an attention temperature correction - Perplexity on long text does not prove the context is being used. You need needle-in-a-haystack (hide a fact in a long text and ask about it) and RULER (a set of synthetic long-context tasks); "Lost in the Middle": retrieval is worse from the middle
Code → nanolm/vit.py: a minimal ViT encoder: patchify, a CLS token (an extra token
whose output serves as the representation of the whole image),
learned positional embeddings, bidirectional attention. Check:
pytest tests/test_vit.py -v; the key test is test_linear_patch_embedding_equals_strided_conv:
a linear layer over flattened patches and a Conv2d with stride = P give the same result.
scripts/rope_extrapolation.py: train on windows of length S, measure perplexity
at S and 2S separately for positions < S and ≥ S, then repeat with base Θ ×4 and ×8.
The mean perplexity over the window hides the gap; you only see it per position.
Math (Track D): D31: contrastive loss and temperature; D32: coupon collector.
Interview question of the week: "The model was trained at 4k, and you need 32k context. What do you do?"
Structure: diagnosis (slow RoPE frequencies are out of distribution) → options
(PI, increasing Θ, YaRN) → short fine-tuning on long documents →
cost (KV cache ×8; attention: FlashAttention, GQA) → validation not by perplexity
but by retrieval across positions.
Sources: Dosovitskiy et al., ViT (2020); Radford et al., CLIP (2021); Liu et al., LLaVA (2023); Chen et al., Positional Interpolation (2023); Peng et al., YaRN (2023); Liu et al., Lost in the Middle (2023); Kingma & Welling, Auto-Encoding Variational Bayes (2013); Goodfellow et al., GAN (2014); Ho et al., DDPM (2020); Rombach et al., Latent Diffusion (2022); Lipman et al., Flow Matching (2023); Sahoo et al., MDLM (2024); Nie et al., LLaDA (2025).
Deeper: 05-ГЛУБИНА, sections "★★ Weeks 16 and 20. Generative models beyond autoregression" and "Small additions".
Week outcomes
- I can implement a ViT encoder from scratch and compute the number of tokens and the cost of attention for a given resolution.
- I can write down the CLIP loss and explain why the batch size equals the number of negatives.
- I can compare a projector, cross-attention and a Q-Former on token count and cost in 3 minutes.
- I can explain which RoPE frequencies break under extrapolation and show it by measuring perplexity per position.
- I can compare autoregression, VAE, diffusion, flow matching, GAN and EBM: what each family models and how it samples.
Self-check
- Why does ViT's cost depend quadratically on resolution rather than linearly?
- How does position interpolation differ from increasing
Θ, and what does each approach lose? - Why does good perplexity at 32k not yet mean the model uses the context?
- Why does an LLM give the exact likelihood of a string while diffusion gives only a lower bound? What does the LLM pay for exactness?
Mock interviews of the week (3 and 4 of 12). (3) ML coding, session B: generation with a KV cache, checked against the full forward pass. (4) Experiment design, session C: "Does the model actually use context beyond 32k, or does it merely not break on it?"