Make sure you check the syllabus for the due date. Please use the notations adopted in class.
Where this HW sits. HW5's networks were stacks of fully-connected layers that ignore the structure of their input. HW6 adds the two architectures built around structure. Convolutional networks (Problems 1–2) exploit the fact that images are local and that the same pattern can appear anywhere. Recurrent networks (Problems 3–4) exploit the fact that sequences unfold in order, carrying a hidden state from one step to the next. Each new layer type is first built from scratch and checked against PyTorch's built-in, then the built-ins are used to go deeper. Problem 4's translation model is the baseline that HW7's attention models and Transformer will be compared against.
Setup. Same environment as HW5; see the NN setup page. A GPU (Apple Silicon or NVIDIA, or Colab's free T4) is recommended for this HW. The starter's default settings keep each training run to roughly 10–15 minutes on a laptop GPU, and its FAST setting (smaller data/model) makes every problem feasible on a CPU. Start early enough that training time isn't your bottleneck.
Instructions. Submit code (Jupyter notebook encouraged) plus a short report with tables, training curves, and example outputs. Libraries (torch, torchvision, d2l, numpy, matplotlib, sacrebleu) are allowed for data handling, plotting, and standard layers. In this HW, "from scratch" means built from torch tensor operations, with autograd allowed (you no longer hand-derive backward passes as in HW5 Problems 1–2), but without the built-in layer being replaced: no nn.Conv2d/F.conv2d inside your convolution, no nn.BatchNorm2d inside your batch norm, no nn.RNN/nn.RNNCell inside your RNN.
Requirements.
Datasets. Fashion-MNIST (28×28 grayscale, 10 classes; auto-download via d2l.FashionMNIST; Problem 1), the same data as HW5, so HW5's MLP is your baseline. CIFAR-10 (32×32 color natural images, 10 classes, 50,000 train / 10,000 test; auto-download via torchvision.datasets.CIFAR10; Problem 2). Penn Treebank (PTB) (about 900k words of Wall Street Journal text, 10k-word vocabulary; auto-download via d2l; Problem 3), the standard word-level language-modeling benchmark, also used by HW7's word vectors. Spanish–English sentence pairs (Tatoeba project, sentence_pairs_large.tsv in the course data folder, ~270k pairs; Problem 4), the same data HW7 uses.
Starter code. The starter notebook provides the device cell, data loading, a shared train/evaluate loop (d2l-style), plotting, and a checkpoint cell after each piece you implement. The checkpoint compares your from-scratch layer against the PyTorch built-in on the same inputs and weights, so you know a layer is right before you train with it. Problem 4's data pipeline and evaluation live in a small module, seq2seq_common.py, that HW7 imports unchanged, so both HWs tokenize, split, decode, and score translations the same way. Answers to the (THEORY) problems, and to written parts inside other problems, may be written directly in the notebook (e.g. a markdown cell) rather than a separate document, as long as they include any equation, plot, or illustration the answer depends on — not prose alone.
See d2l 7.2, 7.4, 7.5, 7.6, and 8.5.
(A — 20 points) Convolution and pooling, from scratch. Build up in the book's order, with a checkpoint (torch.testing.assert_close against F.conv2d) after each step:
(B — 15 points) LeNet, twice. Build the classic LeNet-5 (sigmoid activations, average pooling) from your layers and train it on Fashion-MNIST; then build it from nn.Conv2d/nn.AvgPool2d and train with the same settings. The accuracies should match within noise; report both, and each one's time per epoch. Then add HW5 Problem 2's MLP to the table, with parameter counts. In 2–3 sentences: how does a network with far fewer parameters match or beat the MLP? Name the two assumptions that convolution builds in.
(C — 10 points) Modernize. Switch to ReLU activations and max pooling (library layers from here on) and retrain. Report the change in accuracy and in how quickly training loss falls. Recall HW5 Problem 2(F)'s gradient-norm plot.
(D — 25 points) Batch normalization, from scratch. Implement batch_norm(X, gamma, beta, moving_mean, moving_var, eps, momentum) and a BatchNormScratch(nn.Module) for both fully-connected inputs (statistics per feature, over the batch) and convolutional inputs (statistics per channel, over batch, height, and width):
Checkpoint: the starter compares your module against nn.BatchNorm2d in both modes, after several training-mode forward passes (so the running estimates are tested too). To match exactly, follow PyTorch's convention: the biased variance normalizes the batch, and the unbiased variance updates the running estimate. Insert batch norm after each convolutional and hidden linear layer (BN-LeNet), and train it next to (C)'s network at a larger learning rate than (C) tolerates. Plot both training curves and comment on what batch norm bought you.
Library calls: torch.nn.functional.conv2d, torch.nn.Conv2d, torch.nn.functional.max_pool2d, torch.nn.functional.avg_pool2d, torch.nn.BatchNorm2d, torch.nn.BatchNorm1d.
Same training/evaluation loop as Problem 1, new data: CIFAR-10's color photos are far more varied than Fashion-MNIST's centered, gray, uniform-background items. The problem is about the three practical tools that make CNNs work on natural images: depth (made trainable by residual connections), data augmentation, and pretrained models. See d2l 8.6, 14.1, and 14.2.
(A — 30 points) From LeNet to a residual network. First, as a baseline, adapt Problem 1(D)'s BN-LeNet to 3 input channels and 32×32 images, train it on CIFAR-10, and compare its accuracy with its Fashion-MNIST accuracy. Then implement Residual(nn.Module): two 3×3 convolutions, each followed by batch norm, with ReLU in between, and the input added back before the final ReLU (the skip connection). When the block changes the number of channels or downsamples with stride 2, the skip path needs a 1×1 convolution to match shapes; implement that case too. Checkpoint: output shapes for both block types. Assemble a small ResNet from your blocks using the starter's layout (stages of blocks with increasing channels), train it, and report train/test accuracy and training curves against the BN-LeNet baseline. In 2–3 sentences, explain why the skip connection helps gradients reach early layers (think of the derivative of x + f(x), and of HW5 Problem 2(F)).
(B — 15 points) Data augmentation. Add random cropping (pad by 4, crop back to 32×32) and random horizontal flips to the training data only, and retrain (A)'s ResNet with the same budget. Plot train and test accuracy vs. epoch for both runs and compare the train/test gaps. Which image changes does this augmentation teach the network to ignore, and which does it not (Problem 6 builds on this)? Why must the test set not be augmented?
(C — 15 points) Transfer learning. Take torchvision.models.resnet18 with ImageNet-pretrained weights, replace its final fully-connected layer with a new 10-class layer, and train on a small CIFAR-10 subset (the starter uses 500 images per class, resized and normalized the way the pretrained model expects). Compare three runs on that same subset:
Report a table of test accuracies. In 2–3 sentences: what did ImageNet pretraining transfer, and why does it matter most when labeled data is scarce?
Library calls: torchvision.datasets.CIFAR10, torchvision.transforms.v2.RandomCrop, torchvision.transforms.v2.RandomHorizontalFlip, torchvision.models.resnet18.
A language model predicts the next word from the words so far. It is the simplest task that forces a network to carry information across time, and the same next-token objective trains today's large language models. Quality is measured by perplexity: exp(average cross-entropy per word), roughly the number of words the model is "choosing between" at each step (lower is better). See d2l 9.3, 9.5, 9.6, 9.7, 10.1, and 10.2. (The book works at the character level on a small novel; we use words on PTB, the classic word-level benchmark, so the model learns word embeddings and connects directly to HW7's word vectors and Transformers.)
(A — 30 points) RNN from scratch. The starter loads PTB, builds the vocabulary, and cuts the text into minibatches of input sequences X (length 35) with targets Y = X shifted by one word. First compute a number to beat: the validation perplexity of the unigram baseline, which predicts every word with its training-set frequency and ignores context. Then implement RNNScratch(nn.Module) with parameters Wxh, Whh, bh and an explicit loop over time steps: Ht = tanh(XtWxh + Ht−1Whh + bh). Put it between an nn.Embedding input layer and a linear output layer over the vocabulary. Also implement gradient clipping by global norm yourself: if the norm of all gradients together exceeds θ, rescale them to norm θ. Checkpoint: copy your weights into nn.RNN and assert identical outputs (nn.RNN has two bias vectors; set one of them to zero). Train, and report validation perplexity vs. epoch against the unigram baseline. Then generate text: feed a prompt (e.g. "the company said"), then repeatedly sample the next word from the softmax, at two temperatures.
(B — 15 points) RNN, GRU, LSTM with PyTorch. Swap your cell for nn.RNN, nn.GRU, and nn.LSTM with the same embedding size, hidden size, and training budget (the only code change is the recurrent layer; the LSTM's state is a pair (H, C)). Report a table of validation perplexity and time per epoch for all four models (including yours), and a sample from the best one.
(C — 15 points) How far back do gradients reach? For your trained RNN and LSTM, take a batch of long sequences (e.g. 100 words), compute the loss at the last step only, and measure the gradient norm ||∂LT / ∂Ht|| with respect to the hidden state at every earlier step t (use retain_grad() on the hidden states). Plot it against the distance T − t, log scale, both models on one figure. Problem 5 asks you to explain the difference.
Library calls: torch.nn.RNN, torch.nn.GRU, torch.nn.LSTM, torch.nn.Embedding, torch.nn.utils.clip_grad_norm_.
An encoder RNN reads the Spanish sentence into a hidden state; a decoder RNN, started from that state, generates the English sentence one word at a time. It is Problem 3's language model, conditioned on a source sentence. See d2l 10.5 and 10.7. seq2seq_common.py provides tokenization, vocabularies (with <pad>, <bos>, <eos>), the train/validation/test split, and the BLEU scorer. The starter keeps sentences short to medium in length, so training is quick but the effect of length is still visible.
(A — 25 points) Encoder-decoder with GRUs. Implement:
(B — 10 points) LSTM version. Replace the GRUs with nn.LSTMs. What has to change in how the state is passed from encoder to decoder? Train with the same budget.
(C — 15 points) Evaluate, and save the baseline for HW7. For both models, report test-set BLEU (sacrebleu) overall and by source-sentence length (the starter's length buckets; plot BLEU vs. length for both models), and show ten test translations: a few good ones and a few bad ones. Where does quality fall off, and why might squeezing the entire source sentence into one fixed-size state be the bottleneck? Finally, run the starter's save cell. It writes results_hw6.json (your BLEU numbers and test-set translations) and a model checkpoint. HW7 loads these to compare its attention models and Transformer against your LSTM (a reference file is provided if your run is weak).
Library calls: torch.nn.GRU, torch.nn.LSTM, torch.nn.Embedding, torch.nn.CrossEntropyLoss (with ignore_index), sacrebleu.corpus_bleu.
Your Problem 3(C) plot most likely shows the plain RNN's gradient shrinking roughly exponentially with distance, while the LSTM's decays much more slowly.
In a short written answer, explain in terms of backpropagation through time why the plain RNN's gradient behaves this way (write ∂HT/∂Ht as a product of per-step Jacobians), and what specific mechanism in the LSTM changes that product. Why does gradient clipping, which you implemented in Problem 3(A), fix exploding gradients but not vanishing ones?
A CNN is trained only on images where the object is centered and upright, and performs very well on a held-out test set from the same distribution. But when the same trained model is evaluated on images where the object has been shifted to a corner of the frame, or rotated 90 degrees, accuracy drops sharply.
In a short written answer, explain what property of convolution + pooling gives CNNs some built-in robustness to shifts (distinguish equivariance of convolution from the local invariance of pooling), and why it does not fully protect against the changes described here. Then use your Problem 2(B) results: which of the two failures would your crop-and-flip augmentation help with, and what would you add for the other?
Build two deep networks of equal depth (20+ convolutional layers): one from your Problem 2(A) residual blocks, one identical but with the skip connections removed. Train both with the same budget and plot training loss. He et al. found the deeper plain network ends with higher training error, which is an optimization failure, not overfitting. Do you see it?
Implement an LSTM cell (input, forget, output gates, candidate memory, memory cell) following d2l 10.1, check it against nn.LSTM with copied weights (mind PyTorch's gate ordering in its weight matrices), and use it in Problem 3.
Replace Problem 4's greedy decoding with beam search (d2l 10.8), with a length normalization. Report BLEU for beam sizes 1 (= greedy), 3, 5.
Replace nn.GRU in Problem 4's encoder and decoder with your Problem 3(A) RNNScratch (or your Problem 8 LSTM). How does BLEU compare with the gated cells, especially on long sentences?