CS6140 Machine Learning — Fall 2026

HW7 — Word Vectors, Attention, Transformers

Make sure you check the syllabus for the due date. Please use the notations adopted in class.

Where this HW sits. HW6 ended on two limitations of recurrent networks. The encoder must squeeze a whole sentence into one fixed-size state (your BLEU-vs-length plot), and information has to survive many steps of recurrence (your gradient-vs-distance plot). Attention removes the first limitation by letting every output position look directly at every input position. Transformers then drop recurrence entirely and use attention for everything. This HW builds that path one mechanism at a time:

Every translation model is scored on the same test set as your HW6 LSTM, so the comparison is direct.

Setup. Same environment as HW5–HW6; see the NN setup page. Problem 1 also uses gensim (pip install gensim, if you haven't already). A GPU (Apple Silicon, NVIDIA, or Colab) is recommended for Problems 3 and 4. The starter's defaults keep each training run to roughly 10–15 minutes on a laptop GPU, and its FAST setting makes every problem feasible on a CPU.

Instructions. Submit code (Jupyter notebook encouraged) plus a short report with tables, attention-weight plots, example translations, and training curves. Libraries (torch, d2l, numpy, matplotlib, sacrebleu, gensim) are allowed for data handling, plotting, and standard layers. As in HW6, "from scratch" means built from torch tensor operations (autograd allowed) without the built-in being replaced. Your attention may not call F.scaled_dot_product_attention, your multi-head attention may not call nn.MultiheadAttention, and your Transformer blocks may not use nn.Transformer* classes. Those built-ins appear only in checkpoint cells, as the reference your code is checked against. (nn.Linear, nn.Embedding, nn.LayerNorm, nn.GRU are fine.)

Requirements.

Datasets. Penn Treebank (PTB) (auto-download via d2l; Problem 1), the same corpus as HW6 Problem 3. Spanish–English sentence pairs (sentence_pairs_large.tsv; Problems 3 and 4), loaded through HW6's seq2seq_common.py with the same tokenization, vocabularies, and train/validation/test split. Your HW6 results (results_hw6.json + checkpoint from HW6 Problem 4(C)); if your HW6 translator was weak or you lost the file, use the provided results_hw6_reference.json and say so in your report. Optional problems use the course data folder's Text8 (text8/), 20 Newsgroups (20newsgroup/), and pretrained GloVe vectors (glove.6B.zip).

Starter code. The starter notebook provides the device cell, data loading, HW6's encoder and training/decoding loop (imported from seq2seq_common.py), plotting helpers for attention heatmaps and BLEU-by-length, and a checkpoint cell after each piece you implement. Each checkpoint compares your module against the PyTorch built-in on the same inputs and weights, or tests a property the piece must have (weights sum to 1, masked positions get zero weight, a future token cannot change a past output). Answers to the (THEORY) problems, and to written parts inside other problems, may be written directly in the notebook (e.g. a markdown cell) rather than a separate document, as long as they include any equation, plot, or illustration the answer depends on — not prose alone.


PROBLEM 1 — Word Vectors: Skip-gram with Negative Sampling (PTB)   [50 points]

Every model in this HW (and HW6's) starts with an nn.Embedding that turns each word into a vector, learned as a side effect of the task. Word2vec learns those vectors as the task itself: a word's vector should predict the words around it, so words used in similar contexts end up close together. See d2l 15.1, 15.2, 15.3, 15.4, and the word2vec tutorial.

(A — 15 points) The training data. Implement the three steps that turn raw text into skip-gram training examples:

  1. Subsampling: discard each occurrence of word w with probability 1 − √(t / f(w)), where f(w) is w's relative frequency (t = 10−4), so very frequent words like "the" stop dominating.
  2. Center/context pairs: for each center word, take the words within a random window of size 1 to 5 on each side.
  3. Negative sampling: for each (center, context) pair draw K = 5 "noise" words from the unigram distribution raised to the 3/4 power.

Checkpoint: the starter reports how often "the" survives subsampling, and compares the empirical frequencies of your negative samples with the target distribution.

(B — 15 points) Model and loss. Two nn.Embedding tables (center vectors v, context vectors u), score u·v, and the negative-sampling loss, written by you:

L = −log σ(uo·vc) − ∑k=1K log σ(−uk·vc)

Use a numerically stable log-sigmoid, and mask padded positions. This is HW2's logistic regression: "is this a real context word, or noise?" Train, and plot the loss.

(C — 10 points) Evaluate. For several words (e.g. chip, baby, beautiful, money, president, stock, tuesday), list the 10 nearest neighbors by cosine similarity. Try a few analogies a : b :: c : ? (answer = nearest word to vb − va + vc; PTB is small, so expect partial success). Project the vectors of the 200 most frequent words to 2-D with PCA (as in HW3) and plot them with labels. Do you see clusters (numbers, months, verbs)?

(D — 10 points) Library baseline. Train gensim.models.Word2Vec (skip-gram, negative sampling, same dimension and window) on the same PTB sentences. For the same query words, report the overlap between gensim's top-10 neighbors and yours, and comment on any systematic differences.

Library calls: torch.nn.Embedding, torch.nn.functional.logsigmoid, gensim.models.Word2Vec.


PROBLEM 2 — Attention Mechanisms, from Scratch   [70 points]

Attention is a soft lookup. A query is compared against every key, the scores become weights through a softmax, and the output is the weighted average of the values. This problem builds every attention component that Problems 3–4 use, each followed by a checkpoint, before any of them is trained on real data. See d2l 11.3, 11.5, and 11.6.

(A — 20 points) Scaled dot-product attention, with masking. Batches contain sentences of different lengths, padded to a common length, and a query must never attend to padding.

  1. Implement masked_softmax(scores, valid_lens): set scores beyond each sequence's valid length to a large negative number before the softmax. Checkpoint: every row of weights sums to 1, and masked positions get (numerically) zero weight.
  2. Implement DotProductAttention(nn.Module): softmax(QKT/√d) V with your masked softmax, keeping the weights for plotting. Checkpoint: compare with F.scaled_dot_product_attention on the same Q, K, V and mask. The two use opposite boolean-mask conventions: in F.scaled_dot_product_attention, True means "may attend." A mismatch here is the usual cause of a failing checkpoint.

(B — 10 points) Why divide by √d? Draw queries and keys with independent unit-variance entries, for dimension d ∈ {4, 16, 64, 256, 1024}. For each d, measure the average largest attention weight per query (1 would mean fully one-hot), with and without the 1/√d scaling, and plot both curves. In 2–3 sentences: what is the variance of q·k, what happens to the softmax as d grows without the scaling, and what would that do to the gradients during training?

(C — 20 points) Multi-head attention. Implement MultiHeadAttention(nn.Module):

Checkpoint: copy your weights into nn.MultiheadAttention (batch_first=True; it packs Wq, Wk, Wv into one in_proj_weight) and assert identical outputs with a padding mask.

(D — 20 points) Self-attention: positions and causality. In self-attention the queries, keys, and values all come from the same sequence. Two things must be added before it can model language:

  1. Word order. Show numerically that self-attention alone is permutation-equivariant: shuffling the input tokens shuffles the outputs the same way, so word order is invisible. Then implement the sinusoidal positional encoding, add it to the inputs, and show the property no longer holds. Plot a few of the encoding's dimensions against position.
  2. No peeking. A decoder generating word i must not see words i+1, i+2, …. Add a causal-mask option, so position i attends only to positions ≤ i. Checkpoint: changing a later token leaves all earlier outputs unchanged. Why does this mask let a Transformer decoder train on all target positions in parallel, when HW6's RNN decoder had to step through them one by one?

Library calls: torch.nn.functional.scaled_dot_product_attention, torch.nn.MultiheadAttention (checkpoints only).


PROBLEM 3 — Translation with Attention: Spanish → English   [50 points]

Keep HW6 Problem 4's encoder and training loop unchanged. Change only the decoder, so that at every step it attends over all encoder outputs instead of relying only on the final state. This is the Bahdanau-style attention decoder of d2l 11.4, with Problem 2(A)'s attention as its scoring function.

(A — 25 points) The attention decoder. Implement Seq2SeqAttentionDecoder. At each decoding step, the decoder's previous hidden state is the query, and the encoder outputs at all source positions are the keys and values (masked by the source lengths). The resulting context vector is concatenated with the embedded input word and fed into the GRU. Train with the same budget as HW6.

(B — 10 points) Look at the alignments. For 3–4 test sentences, plot the attention weights as a heatmap (target words × source words). Is the alignment roughly diagonal? Find a case where it isn't, e.g. a Spanish adjective that follows its noun (la casa blanca → the white house), and explain what the model is doing there.

(C — 15 points) Compare with HW6. Report test BLEU overall and by source length for this model next to your HW6 GRU and LSTM (from results_hw6.json): one table, one BLEU-vs-length plot with all three. Which length bucket gains the most, and why is that what attention should do?

Library calls: none beyond Problem 2's; the HW6 LSTM is the baseline.


PROBLEM 4 — The Transformer: Spanish → English   [60 points]

Remove recurrence entirely ("attention is all you need"). See d2l 11.7.

(A — 30 points) Build it from your parts. Using your Problem 2 multi-head attention, positional encoding, and causal mask, implement:

Checkpoint: output shapes, and the causal property from Problem 2(D), for the whole decoder.

(B — 15 points) Train and compare. Train with the starter's settings (Adam with a learning-rate warm-up schedule; the settings are provided because Transformers are sensitive to them). Report test BLEU overall and by length next to the HW6 LSTM and Problem 3's attention GRU. Give one table, including parameter counts and time per epoch, and one BLEU-vs-length plot with all models. Greedy decoding still runs one word at a time at test time. Only training is parallel.

(C — 15 points) Analysis. For one test sentence, plot the decoder's cross-attention weights (last layer, each head separately) and compare them with Problem 3's single alignment. Plot one encoder self-attention layer's heads too. Then give an honest comparison in 3–4 sentences: on this dataset size and training budget, where does the Transformer win, and where doesn't it? What would you expect to change with 100× more data?

Library calls: torch.nn.LayerNorm; optionally torch.nn.Transformer with the same sizes, as a library check on your BLEU.


PROBLEM 5 (THEORY) — Why Multiple Attention Heads Instead of One Larger Head   [25 points]

Suppose you replace your multi-head attention with a single head of the same total dimensionality (so the number of parameters is roughly unchanged), keeping everything else fixed. Performance drops on a task that requires tracking several different relationships between words at once, e.g. syntactic ones such as subject–verb agreement, and coreference between a pronoun and the noun it refers to.

In a short written answer, explain conceptually why several heads can help here: what can h separate softmaxes express that one softmax over the same dimensions cannot? Use your Problem 4(C) per-head plots as evidence: do different heads attend differently?


PROBLEM 6 (THEORY) — Self-Attention vs. Cross-Attention   [25 points]

Your Problem 4 Transformer uses three kinds of attention: encoder self-attention, decoder (causal) self-attention, and decoder-to-encoder cross-attention. Explain the structural difference between self-attention and cross-attention in terms of where Q, K, and V come from. Explain why translation needs both, and what each would be unable to do without the other. Which of the three is the only one that needs a causal mask, and why?



OPTIONAL PROBLEMS [no credit]

PROBLEM 7 — A Tiny GPT: Transformer Language Model on PTB   [optional, no credit; GPU recommended]

Drop the encoder and the cross-attention from your Problem 4 decoder, keeping causal self-attention and the FFN. What remains is a decoder-only Transformer, the architecture of GPT-style language models. Train it on HW6 Problem 3's exact PTB setup (same vocabulary, sequence length, and budget) and compare its validation perplexity with your HW6 RNN, GRU, and LSTM. Sample text from it with the same prompts.


PROBLEM 8 — Pretrained Transformers: Where Your Models Fit   [optional, no credit]

Read d2l 11.9. Write a short comparison: which of your HW7 models corresponds to an encoder-only model (BERT), a decoder-only model (GPT), and an encoder-decoder model (T5)? What pretraining objective does each use, and how does it relate to the objectives you trained on (skip-gram in Problem 1, next-word prediction in HW6 Problem 3, translation in Problem 4)?


PROBLEM 9 — Self-Attention Inside an RNN   [optional, no credit]

Hybrid models from previous years' HW7: add self-attention inside the GRU encoder (and then the decoder) of Problem 3's model, keeping its cross-attention. Where does the hybrid land between Problem 3's model and the full Transformer in BLEU?


PROBLEM 10 — Word Vectors on Larger Corpora, and Pretrained GloVe   [optional, no credit]

Train Problem 1's model on Text8 (Wikipedia, ~17M words) or 20 Newsgroups, and compare neighbor lists with PTB's for words such as China, computer, phone, God, Napoleon, Catholic. Then load pretrained GloVe vectors (glove.6B.zip) and compare again. Finally, initialize Problem 3's embedding layers with GloVe vectors. Does BLEU change?


References