{ "cells": [ { "cell_type": "markdown", "id": "cd95d2c4", "metadata": {}, "source": [ "# CS6140 Machine Learning — Fall 2026\n", "# HW5 Starter — Neural Networks: Backprop, Training, VAE, GAN\n", "\n", "This notebook implements `HW5_26F.html`. Read the HTML for the full problem statements; the markdown\n", "cells here repeat the key instructions next to the code.\n", "\n", "| Problem | Topic | Points |\n", "|---|---|---|\n", "| 1 | Backpropagation from scratch: the 8-3-8 autoencoder | 50 |\n", "| 2 | A classifier network: build it (numpy, then PyTorch/d2l), then train it well | 100 |\n", "| 3 | Variational autoencoder (VAE) on MNIST | 50 |\n", "| 4 | Generative adversarial network (GAN): run it and explain it (code given) | 10 |\n", "| 5, 6 | THEORY: dropout; diagnosing the learning rate | 25 + 25 |\n", "| 8 | optional: backprop in matrix form (7, DCGAN, is HTML-only) | no credit |\n", "\n", "**How to work through this notebook.** Everything runs as given except the lines marked\n", "`### TODO_STUDENT ###`: their code has been replaced by `...` plus a one-line description of what\n", "goes there. Replace each `...` with real code. Until you do, the cell (or a later cell that uses it)\n", "will fail or print nonsense. That's expected, not a bug. Written answers go in the markdown cells\n", "that say *Your answer here*.\n", "\n", "**Checkpoint cells** follow each piece you implement. They test your code against a numerical\n", "check or the matching PyTorch function, and fail with a message if something is off. Don't move\n", "on until the checkpoint above you passes: a network trained on a wrong gradient can still *look*\n", "like it's learning.\n", "\n", "**Work top to bottom.** Problem 2 reuses Problem 1's layer classes; Problem 3's plain autoencoder\n", "reuses your VAE; Problem 4's GAN is complete code to run and explain. The whole notebook, once filled in,\n", "runs in about 3–5 minutes on a laptop CPU.\n", "\n", "**Setup.** If the first cell fails, see the course's *NN setup page* (`html/NN_setup_26F.html`).\n", "Datasets are read from the course `data/` folder, two levels up (`../../data`); Fashion-MNIST and\n", "MNIST are downloaded into `../../data/torchvision/` on first run (~150 MB).\n", "\n", "**Requirements.**\n", "1. Fill in every `TODO_STUDENT` blank in the required problems according to the algorithm steps\n", " covered in lecture and the linked materials, and make sure the notebook runs top to bottom\n", " without errors. (Blanks in the optional problems are not required.)\n", "2. Understand the *whole* notebook — including the provided/given code, not just the blanks you\n", " filled in — well enough to explain any part of it during office hours.\n" ] }, { "cell_type": "code", "execution_count": null, "id": "bf4a0b93", "metadata": {}, "outputs": [], "source": [ "import sys\n", "import time\n", "import numpy as np\n", "import matplotlib.pyplot as plt\n", "import torch\n", "from torch import nn\n", "import torch.nn.functional as F\n", "import torchvision\n", "from d2l import torch as d2l\n", "from tqdm import tqdm\n", "from IPython.display import display\n", "\n", "\n", "def get_device():\n", " if torch.cuda.is_available(): # NVIDIA GPU\n", " return torch.device('cuda')\n", " if torch.backends.mps.is_available(): # Apple Silicon GPU\n", " return torch.device('mps')\n", " return torch.device('cpu')\n", "\n", "\n", "DEVICE = get_device()\n", "if DEVICE.type == 'mps': # d2l only knows CUDA: point its GPU helpers at Apple's GPU\n", " d2l.num_gpus = lambda: 1\n", " d2l.gpu = lambda i=0: DEVICE\n", " d2l.try_gpu = lambda i=0: DEVICE\n", "print('Best available device:', DEVICE)\n", "\n", "# HW5's networks are small. On the machines we tested, the CPU trains them as fast as, or faster\n", "# than, a GPU (moving tiny batches to the GPU costs more than the math), so HW5 trains everything on\n", "# the CPU. HW6-HW7's bigger models will use DEVICE.\n", "TRAIN_DEVICE = torch.device('cpu')\n", "\n", "SEED = 0\n", "np.random.seed(SEED)\n", "torch.manual_seed(SEED)\n", "\n", "DATA = \"../../data\" # the course data folder\n", "TORCH_DATA = f\"{DATA}/torchvision\" # Fashion-MNIST + MNIST download here (~150 MB)\n", "GAUSS = \"../../3_generative_models/HW3\" # HW4's 2-D Gaussian-mixture files\n", "\n", "\n", "# --- Progress bars (nothing to do here) ---\n", "class _BarOutput:\n", " \"\"\"Where tqdm writes its bar: ONE notebook output, updated in place (not a new line per redraw).\n", " Inside a d2l live plot, the bar becomes the plot's title instead, because d2l clears the cell's\n", " output on every redraw and would erase a separate bar.\"\"\"\n", "\n", " def __init__(self, board=None):\n", " self.board, self.handle = board, None\n", "\n", " def write(self, text):\n", " text = text.strip(\"\\r\\n\")\n", " if not text:\n", " return\n", " if self.board is not None:\n", " if self.board.fig is not None:\n", " self.board.fig.suptitle(text, fontsize=8, family=\"monospace\")\n", " elif self.handle is None:\n", " self.handle = display({\"text/plain\": text}, raw=True, display_id=True)\n", " else:\n", " self.handle.update({\"text/plain\": text}, raw=True)\n", "\n", " def flush(self):\n", " pass\n", "\n", "\n", "def progress(total, desc, unit=\"batch\", board=None):\n", " \"\"\"A progress bar for a training loop: call .update(1) per step, .close() at the end.\"\"\"\n", " width = 25 if board is None else 12\n", " fmt = \"{desc}: {percentage:3.0f}%|{bar:\" + str(width) + \"}| {n_fmt}/{total_fmt} [{elapsed}<{remaining}, {rate_fmt}{postfix}]\"\n", " return tqdm(total=total, desc=desc, unit=unit, file=_BarOutput(board), mininterval=0.5, ascii=False,\n", " bar_format=fmt)\n", "\n", "\n", "def _fit_with_progress(self, model, data):\n", " \"\"\"d2l.Trainer.fit, plus a progress bar over all training batches (shown as the live plot's title).\"\"\"\n", " self.prepare_data(data)\n", " self.prepare_model(model)\n", " self.optim = model.configure_optimizers()\n", " self.epoch = self.train_batch_idx = self.val_batch_idx = 0\n", " board = model.board if model.board.display else None\n", " if board is not None:\n", " board.figsize = (5.5, 3.2) # room for the bar in the title\n", " bar = progress(self.max_epochs * self.num_train_batches, type(model).__name__, board=board)\n", " loader = self.train_dataloader\n", " self.train_dataloader = _TickingLoader(loader, bar)\n", " try:\n", " for self.epoch in range(self.max_epochs):\n", " self.fit_epoch()\n", " bar.set_postfix(epoch=self.epoch + 1)\n", " finally:\n", " self.train_dataloader = loader\n", " bar.close()\n", "\n", "\n", "class _TickingLoader:\n", " \"\"\"Wraps a DataLoader so that every batch it hands out advances the progress bar.\"\"\"\n", " def __init__(self, loader, bar):\n", " self.loader, self.bar = loader, bar\n", " def __iter__(self):\n", " for batch in self.loader:\n", " yield batch\n", " self.bar.update(1)\n", " def __len__(self):\n", " return len(self.loader)\n", "\n", "\n", "d2l.Trainer.fit = _fit_with_progress\n" ] }, { "cell_type": "markdown", "id": "321e3b2a", "metadata": {}, "source": [ "## Data\n", "\n", "* **Wine** (Problem 2, for debugging): 3 classes, 13 features. Column 0 of the CSV is the class label\n", " (1, 2, 3), which we shift to 0, 1, 2. Features are standardized with *training* statistics.\n", "* **Fashion-MNIST** (Problem 2): 28×28 grayscale clothing images, 10 classes. It's the same data\n", " `d2l.FashionMNIST` wraps; we load it through `torchvision` so that numpy (Part I) and PyTorch\n", " (Parts I–II) see exactly the same arrays. Each image is flattened to a vector of 784 pixel values\n", " in [0, 1].\n", "* **MNIST** (Problem 3): handwritten digits, same format.\n", "* **2gaussian.txt / 3gaussian.txt** (Problems 4, 7): HW4's 2-D Gaussian-mixture data.\n" ] }, { "cell_type": "code", "execution_count": null, "id": "913ea0cc", "metadata": {}, "outputs": [], "source": [ "def load_wine(path):\n", " a = np.loadtxt(path, delimiter=\",\")\n", " return a[:, 1:], a[:, 0].astype(int) - 1 # column 0 = class label 1..3 -> 0..2\n", "\n", "\n", "X_wine_tr, y_wine_tr = load_wine(f\"{DATA}/train_wine.csv\")\n", "X_wine_te, y_wine_te = load_wine(f\"{DATA}/test_wine.csv\")\n", "wine_mean, wine_std = X_wine_tr.mean(axis=0), X_wine_tr.std(axis=0)\n", "X_wine_tr = (X_wine_tr - wine_mean) / wine_std\n", "X_wine_te = (X_wine_te - wine_mean) / wine_std\n", "print(\"Wine:\", X_wine_tr.shape, X_wine_te.shape, \"class counts (train):\", np.bincount(y_wine_tr))\n", "\n", "\n", "def load_images(dataset_class, train):\n", " \"\"\"Flattened images as float64 numpy (pixels in [0,1]) and integer labels.\"\"\"\n", " ds = dataset_class(root=TORCH_DATA, train=train, download=True)\n", " X = ds.data.reshape(len(ds), -1).numpy().astype(np.float64) / 255.0\n", " return X, ds.targets.numpy()\n", "\n", "\n", "X_fm_tr, y_fm_tr = load_images(torchvision.datasets.FashionMNIST, train=True)\n", "X_fm_te, y_fm_te = load_images(torchvision.datasets.FashionMNIST, train=False)\n", "FM_CLASSES = [\"T-shirt\", \"Trouser\", \"Pullover\", \"Dress\", \"Coat\",\n", " \"Sandal\", \"Shirt\", \"Sneaker\", \"Bag\", \"Ankle boot\"]\n", "print(\"Fashion-MNIST:\", X_fm_tr.shape, X_fm_te.shape)\n", "\n", "# float32 torch copies of the same arrays, for the PyTorch parts\n", "Xt_fm_tr, yt_fm_tr = torch.tensor(X_fm_tr, dtype=torch.float32), torch.tensor(y_fm_tr)\n", "Xt_fm_te, yt_fm_te = torch.tensor(X_fm_te, dtype=torch.float32), torch.tensor(y_fm_te)\n", "\n", "X_mn_tr, y_mn_tr = load_images(torchvision.datasets.MNIST, train=True)\n", "X_mn_te, y_mn_te = load_images(torchvision.datasets.MNIST, train=False)\n", "Xt_mn_tr, yt_mn_tr = torch.tensor(X_mn_tr, dtype=torch.float32), torch.tensor(y_mn_tr)\n", "Xt_mn_te, yt_mn_te = torch.tensor(X_mn_te, dtype=torch.float32), torch.tensor(y_mn_te)\n", "print(\"MNIST:\", X_mn_tr.shape, X_mn_te.shape)\n", "\n", "gauss2 = np.loadtxt(f\"{GAUSS}/2gaussian.txt\")\n", "gauss3 = np.loadtxt(f\"{GAUSS}/3gaussian.txt\")\n", "print(\"2gaussian:\", gauss2.shape, \" 3gaussian:\", gauss3.shape)\n", "\n", "fig, axes = plt.subplots(2, 10, figsize=(12, 2.8))\n", "for i in range(10):\n", " axes[0, i].imshow(X_fm_tr[np.where(y_fm_tr == i)[0][0]].reshape(28, 28), cmap=\"gray\")\n", " axes[0, i].set_title(FM_CLASSES[i], fontsize=8)\n", " axes[1, i].imshow(X_mn_tr[np.where(y_mn_tr == i)[0][0]].reshape(28, 28), cmap=\"gray\")\n", "for ax in axes.ravel():\n", " ax.axis(\"off\")\n", "plt.suptitle(\"Fashion-MNIST (top) and MNIST (bottom)\")\n", "plt.show()\n" ] }, { "cell_type": "markdown", "id": "34afc050", "metadata": {}, "source": [ "---\n", "## Problem 1 — Backpropagation from Scratch: the 8-3-8 Autoencoder [50 points]\n", "\n", "A network with 8 inputs, 3 hidden units, 8 outputs, all sigmoid, trained so that each of the 8\n", "one-hot inputs is reproduced at the output (see the table in the HTML). The point is *how* you build\n", "it: as a chain of layers, each knowing its own forward and backward pass — the same design PyTorch\n", "uses (d2l 5.3).\n", "\n", "**Conventions** (row-vector convention, as in the Regression notes). A batch `x` has shape\n", "(n, n_in). A `Linear` layer computes `x @ W + b` with `W` of shape (n_in, n_out). `backward(grad_out)`\n", "receives dL/d(output) of shape (n, n_out) and returns dL/d(input) of shape (n, n_in). The square\n", "loss is averaged over the batch: L = (1/2n) Σᵢ ‖ŷᵢ − yᵢ‖².\n", "\n", "### (A) Layers with forward and backward [15 points]\n" ] }, { "cell_type": "code", "execution_count": null, "id": "ad4b200d", "metadata": {}, "outputs": [], "source": [ "class Linear:\n", " \"\"\"Fully-connected layer: out = x @ W + b, with x of shape (n, n_in).\"\"\"\n", "\n", " def __init__(self, n_in, n_out, rng, scale=None):\n", " # default init = PyTorch's nn.Linear default: uniform in +-1/sqrt(n_in)\n", " bound = 1.0 / np.sqrt(n_in) if scale is None else scale\n", " self.W = rng.uniform(-bound, bound, size=(n_in, n_out))\n", " self.b = rng.uniform(-bound, bound, size=n_out)\n", " self.dW = np.zeros_like(self.W)\n", " self.db = np.zeros_like(self.b)\n", "\n", " def forward(self, x):\n", " self.x = x # cache the input: backward needs it for dL/dW\n", " return ... ### TODO_STUDENT ### the affine map x W + b\n", "\n", " def backward(self, grad_out):\n", " self.dW = ... ### TODO_STUDENT ### dL/dW = x^T (dL/dout) (sums over the batch)\n", " self.db = ... ### TODO_STUDENT ### dL/db = dL/dout summed over the batch\n", " return ... ### TODO_STUDENT ### dL/dx = (dL/dout) W^T, handed to the previous layer\n", "\n", "\n", "class Sigmoid:\n", " def forward(self, x):\n", " self.out = ... ### TODO_STUDENT ### sigma(x); cache it, backward needs it\n", " return self.out\n", "\n", " def backward(self, grad_out):\n", " return ... ### TODO_STUDENT ### chain rule with sigma'(x) = sigma(x) (1 - sigma(x))\n", "\n", "\n", "class SquareLoss:\n", " \"\"\"L = (1/2n) sum_i ||y_hat_i - y_i||^2.\"\"\"\n", "\n", " def forward(self, y_hat, y):\n", " self.diff = y_hat - y\n", " return ... ### TODO_STUDENT ### the loss value\n", "\n", " def backward(self):\n", " return ... ### TODO_STUDENT ### dL/dy_hat\n", "\n", "\n", "class Network:\n", " \"\"\"A list of layers plus a loss layer. forward runs the layers in order; backward runs them\n", " in reverse, each layer handing its input-gradient to the layer before it.\"\"\"\n", "\n", " def __init__(self, layers, loss):\n", " self.layers, self.loss = layers, loss\n", "\n", " def forward(self, x):\n", " for layer in self.layers:\n", " x = ... ### TODO_STUDENT ### run each layer in order\n", " return x\n", "\n", " def backward(self):\n", " grad = self.loss.backward()\n", " for layer in reversed(self.layers):\n", " grad = ... ### TODO_STUDENT ### pass the gradient back through each layer, last to first\n", "\n", " def linear_layers(self):\n", " return [layer for layer in self.layers if isinstance(layer, Linear)]\n", "\n", " def sgd_step(self, lr):\n", " for layer in self.linear_layers():\n", " layer.W -= ... ### TODO_STUDENT ### gradient descent update of W\n", " layer.b -= ... ### TODO_STUDENT ### gradient descent update of b\n" ] }, { "cell_type": "markdown", "id": "d898e3d4", "metadata": {}, "source": [ "### (B) Check the gradients, two ways [15 points]\n", "\n", "Both checkers are given. Your job is to make them pass.\n", "\n", "1. **Numerically:** every weight's backprop gradient vs. the centered finite difference\n", " (L(w+ε) − L(w−ε)) / 2ε, with ε = 10⁻⁵, in float64. The maximum relative error should be below 10⁻⁶.\n", "2. **Against PyTorch autograd:** the same network built from `nn.Linear` + `nn.Sigmoid`, with *your*\n", " weights copied in; PyTorch's gradients must equal yours.\n" ] }, { "cell_type": "code", "execution_count": null, "id": "e55169dd", "metadata": {}, "outputs": [], "source": [ "def numerical_gradient_check(net, X, Y, eps=1e-5):\n", " \"\"\"Max relative error between backprop and centered finite differences, over ALL weights and biases.\"\"\"\n", " net.loss.forward(net.forward(X), Y)\n", " net.backward()\n", " worst = 0.0\n", " for li, layer in enumerate(net.linear_layers()):\n", " for name, P, dP in [(\"W\", layer.W, layer.dW.copy()), (\"b\", layer.b, layer.db.copy())]:\n", " num = np.zeros_like(P)\n", " for idx in np.ndindex(P.shape):\n", " old = P[idx]\n", " P[idx] = old + eps\n", " L_plus = net.loss.forward(net.forward(X), Y)\n", " P[idx] = old - eps\n", " L_minus = net.loss.forward(net.forward(X), Y)\n", " P[idx] = old\n", " num[idx] = (L_plus - L_minus) / (2 * eps)\n", " rel = np.abs(num - dP) / np.maximum(np.abs(num) + np.abs(dP), 1e-8)\n", " print(f\" Linear layer {li}, {name}: max relative error {rel.max():.2e}\")\n", " worst = max(worst, rel.max())\n", " return worst\n", "\n", "\n", "def autograd_check(net, X, Y):\n", " \"\"\"Rebuild `net` from nn layers with the SAME weights and compare PyTorch's gradients to yours.\"\"\"\n", " torch_layers = []\n", " for layer in net.layers:\n", " if isinstance(layer, Linear):\n", " lin = nn.Linear(*layer.W.shape).double()\n", " with torch.no_grad():\n", " lin.weight.copy_(torch.tensor(layer.W.T)) # nn.Linear stores W transposed: (n_out, n_in)\n", " lin.bias.copy_(torch.tensor(layer.b))\n", " torch_layers.append(lin)\n", " elif type(layer).__name__ == \"Sigmoid\":\n", " torch_layers.append(nn.Sigmoid())\n", " elif type(layer).__name__ == \"ReLU\":\n", " torch_layers.append(nn.ReLU())\n", " tnet = nn.Sequential(*torch_layers)\n", " Xt, Yt = torch.tensor(X), torch.tensor(Y)\n", " loss = 0.5 * nn.MSELoss(reduction=\"sum\")(tnet(Xt), Yt) / len(X) # the same loss as SquareLoss\n", " loss.backward()\n", " net.loss.forward(net.forward(X), Y)\n", " net.backward()\n", " torch_linears = [m for m in tnet if isinstance(m, nn.Linear)]\n", " for lin, layer in zip(torch_linears, net.linear_layers()):\n", " torch.testing.assert_close(lin.weight.grad.T, torch.tensor(layer.dW))\n", " torch.testing.assert_close(lin.bias.grad, torch.tensor(layer.db))\n", " print(\" PyTorch autograd agrees with your backprop on every weight and bias.\")\n", "\n", "\n", "# CHECKPOINT: a fresh 8-3-8 network, before any training\n", "X_ae = np.eye(8) # the 8 one-hot inputs; the targets are the same 8 vectors\n", "rng = np.random.default_rng(SEED)\n", "check_net = Network([Linear(8, 3, rng, scale=0.5), Sigmoid(), Linear(3, 8, rng, scale=0.5), Sigmoid()], SquareLoss())\n", "print(\"(i) numerical gradient check:\")\n", "worst = numerical_gradient_check(check_net, X_ae, X_ae)\n", "assert worst < 1e-6, f\"max relative error {worst:.1e} >= 1e-6: some backward() is wrong (see the per-layer lines above)\"\n", "print(\"(ii) autograd check:\")\n", "autograd_check(check_net, X_ae, X_ae)\n", "print(\"CHECKPOINT PASSED\")\n" ] }, { "cell_type": "markdown", "id": "d4392a5b", "metadata": {}, "source": [ "### (C) Train it, and read what it learned [20 points]\n", "\n", "Full-batch gradient descent on all 8 examples, from random initial weights (uniform in ±0.5, so the\n", "network doesn't start at a solution). The learning rate is large (5.0) because the loss is averaged\n", "over only 8 examples and the sigmoid's slope is at most 1/4; 5,000 iterations take well under a second.\n" ] }, { "cell_type": "code", "execution_count": null, "id": "2865ed02", "metadata": {}, "outputs": [], "source": [ "rng = np.random.default_rng(1)\n", "ae = Network([Linear(8, 3, rng, scale=0.5), Sigmoid(), Linear(3, 8, rng, scale=0.5), Sigmoid()], SquareLoss())\n", "lr, n_iters = 5.0, 5000\n", "losses = []\n", "for it in range(n_iters):\n", " loss = ... ### TODO_STUDENT ### forward pass and loss on all 8 examples\n", " ... ### TODO_STUDENT ### backpropagate\n", " ... ### TODO_STUDENT ### one gradient descent step\n", " losses.append(loss)\n", "\n", "plt.figure(figsize=(5, 3))\n", "plt.semilogy(losses)\n", "plt.xlabel(\"iteration\"); plt.ylabel(\"square loss (log scale)\"); plt.title(\"8-3-8 autoencoder training\")\n", "plt.show()\n", "\n", "ae.forward(X_ae)\n", "H = ae.layers[1].out # the hidden (Sigmoid) layer's output for the 8 inputs\n", "recon = ae.forward(X_ae).argmax(axis=1)\n", "print(\"input hidden values rounded reconstructed as\")\n", "for i in range(8):\n", " print(f\" {i} {np.array2string(H[i], precision=2, floatmode='fixed')} {''.join(str(int(v > 0.5)) for v in H[i])} {recon[i]}\")\n", "codes = {tuple((H[i] > 0.5).astype(int)) for i in range(8)}\n", "print(f\"final loss {losses[-1]:.4f}; distinct 3-bit codes: {len(codes)} of 8; all 8 reconstructed: {np.all(recon == np.arange(8))}\")\n" ] }, { "cell_type": "markdown", "id": "78e0b68a", "metadata": {}, "source": [ "**Written answer (C).** Since the outputs equal the inputs, this network is an encoder-decoder. What does training achieve in this context, and why is the 3-unit hidden layer (rather than, say, 8 units) essential to that?\n", "\n", "*How to think about it.* Round each hidden value in your table to 0 or 1. How many different patterns can 3 binary units show, and how many inputs must be told apart? Then imagine the same network with 8 hidden units: what is the laziest way it could reproduce its input, and would its hidden layer then tell you anything about the data? Think of the hidden layer as the only \"message\" the encoder may send to the decoder.\n", "\n", "*Your answer here.*\n" ] }, { "cell_type": "markdown", "id": "54875e8e", "metadata": {}, "source": [ "---\n", "## Problem 2 — A Classifier Network: Build It, Then Train It Well (Fashion-MNIST) [100 points]\n", "\n", "**Part I** builds a multiclass network three ways: from scratch, with the wrong loss, and in\n", "PyTorch. **Part II** takes the PyTorch version and studies the training choices that decide whether\n", "a network learns well, badly, or not at all.\n", "\n", "### Part I — Build it\n", "\n", "### (A) From scratch: ReLU, softmax + cross-entropy, minibatch SGD [25 points]\n", "\n", "The output layer is a **softmax**, pᵢₖ = exp(oᵢₖ) / Σⱼ exp(oᵢⱼ), trained with **cross-entropy**\n", "L = −(1/n) Σᵢ log pᵢ,yᵢ — maximum likelihood, the same objective as HW2's logistic regression, now\n", "with a hidden layer in front. Written as one combined layer, its backward pass is just\n", "(softmax(o) − onehot(y)) / n: see Equation 4 of the NN_BASICS note (§5.2); the derivation is in\n", "§2 of the NN_ADVANCED note.\n" ] }, { "cell_type": "code", "execution_count": null, "id": "acb5a1fe", "metadata": {}, "outputs": [], "source": [ "class ReLU:\n", " def forward(self, x):\n", " self.mask = ... ### TODO_STUDENT ### remember where the input was positive\n", " return x * self.mask\n", "\n", " def backward(self, grad_out):\n", " return ... ### TODO_STUDENT ### ReLU'(x) = 1 where x > 0, else 0\n", "\n", "\n", "class SoftmaxCrossEntropy:\n", " \"\"\"Softmax output + cross-entropy loss as ONE layer. y holds integer labels 0..K-1.\"\"\"\n", "\n", " def forward(self, o, y):\n", " o = ... ### TODO_STUDENT ### subtract each row's max (numerical stability; softmax is unchanged)\n", " e = np.exp(o)\n", " self.p = ... ### TODO_STUDENT ### softmax probabilities, one row per example\n", " self.y = y\n", " return ... ### TODO_STUDENT ### mean negative log-probability of the true class\n", "\n", " def backward(self):\n", " grad = self.p.copy()\n", " grad[np.arange(len(self.y)), self.y] -= ... ### TODO_STUDENT ### softmax(o) - onehot(y)\n", " return grad / len(self.y)\n", "\n", "\n", "# CHECKPOINT: gradient check of the new layers on a small Wine batch\n", "rng = np.random.default_rng(SEED)\n", "chk = Network([Linear(13, 5, rng), ReLU(), Linear(5, 3, rng)], SoftmaxCrossEntropy())\n", "worst = numerical_gradient_check(chk, X_wine_tr[:20], y_wine_tr[:20])\n", "assert worst < 1e-6, f\"max relative error {worst:.1e}: check ReLU / SoftmaxCrossEntropy backward\"\n", "print(\"CHECKPOINT PASSED\")\n" ] }, { "cell_type": "code", "execution_count": null, "id": "3f542083", "metadata": {}, "outputs": [], "source": [ "def accuracy_scratch(net, X, y):\n", " return np.mean(net.forward(X).argmax(axis=1) == y)\n", "\n", "\n", "def train_scratch(net, X, y, lr, epochs, batch_size, rng, X_test=None, y_test=None, targets=None, desc=\"from scratch\"):\n", " \"\"\"Minibatch SGD. `targets` (optional) is what the loss compares against instead of y -- e.g.\n", " one-hot vectors for the square loss in (B). Accuracy is always measured against the labels y.\"\"\"\n", " T = y if targets is None else targets\n", " hist = {\"loss\": [], \"train_acc\": [], \"test_acc\": []}\n", " bar = progress(epochs * -(-len(X) // batch_size), desc)\n", " for epoch in range(epochs):\n", " perm = rng.permutation(len(X)) # a fresh random order every epoch\n", " for start in range(0, len(X), batch_size):\n", " idx = perm[start:start + batch_size]\n", " loss = ... ### TODO_STUDENT ### forward pass + loss on this minibatch\n", " ... ### TODO_STUDENT ### backpropagate\n", " ... ### TODO_STUDENT ### SGD update\n", " hist[\"loss\"].append(loss)\n", " bar.update(1)\n", " hist[\"train_acc\"].append(accuracy_scratch(net, X, y))\n", " bar.set_postfix(epoch=epoch + 1, train_acc=f\"{hist['train_acc'][-1]:.3f}\")\n", " if X_test is not None:\n", " hist[\"test_acc\"].append(accuracy_scratch(net, X_test, y_test))\n", " bar.close()\n", " return hist\n", "\n", "\n", "# --- Wine: 13-k-3, trains in well under a second ---\n", "print(\"Wine, from scratch (lr=0.1, 100 epochs, batch 16):\")\n", "wine_scratch = {}\n", "for k in [4, 16, 64]:\n", " rng = np.random.default_rng(SEED)\n", " net = Network([Linear(13, k, rng), ReLU(), Linear(k, 3, rng)], SoftmaxCrossEntropy())\n", " train_scratch(net, X_wine_tr, y_wine_tr, lr=0.1, epochs=100, batch_size=16, rng=rng, desc=f\"Wine, k={k}\")\n", " wine_scratch[k] = (accuracy_scratch(net, X_wine_tr, y_wine_tr), accuracy_scratch(net, X_wine_te, y_wine_te))\n", " print(f\" k={k:3d}: train {wine_scratch[k][0]:.3f} test {wine_scratch[k][1]:.3f}\")\n", "\n", "# --- Fashion-MNIST: 784-256-10, the same code ---\n", "rng = np.random.default_rng(SEED)\n", "fm_scratch = Network([Linear(784, 256, rng), ReLU(), Linear(256, 10, rng)], SoftmaxCrossEntropy())\n", "t0 = time.time()\n", "hist_fm_ce = train_scratch(fm_scratch, X_fm_tr, y_fm_tr, lr=0.1, epochs=10, batch_size=256, rng=rng,\n", " X_test=X_fm_te, y_test=y_fm_te, desc=\"Fashion-MNIST, cross-entropy\")\n", "print(f\"\\nFashion-MNIST, from scratch (lr=0.1, 10 epochs, batch 256): \"\n", " f\"train {hist_fm_ce['train_acc'][-1]:.3f} test {hist_fm_ce['test_acc'][-1]:.3f} ({time.time() - t0:.1f} s)\")\n", "\n", "fig, axes = plt.subplots(1, 2, figsize=(10, 3))\n", "axes[0].plot(hist_fm_ce[\"loss\"], alpha=0.4, lw=0.5)\n", "axes[0].plot(np.convolve(hist_fm_ce[\"loss\"], np.ones(50) / 50, mode=\"valid\"), lw=1.5, label=\"moving average\")\n", "axes[0].set_xlabel(\"iteration\"); axes[0].set_ylabel(\"cross-entropy\"); axes[0].legend()\n", "axes[1].plot(range(1, 11), hist_fm_ce[\"train_acc\"], \"o-\", label=\"train\")\n", "axes[1].plot(range(1, 11), hist_fm_ce[\"test_acc\"], \"s-\", label=\"test\")\n", "axes[1].set_xlabel(\"epoch\"); axes[1].set_ylabel(\"accuracy\"); axes[1].legend()\n", "plt.suptitle(\"From-scratch MLP on Fashion-MNIST\"); plt.show()\n" ] }, { "cell_type": "markdown", "id": "2e8da8f9", "metadata": {}, "source": [ "### (B) Why cross-entropy? [10 points]\n", "\n", "The same network and budget (10 epochs, **same learning rate 0.1**), but with Problem 1's recipe:\n", "sigmoid outputs and square loss against one-hot targets. The two losses aren't on the same scale, so\n", "we compare test accuracy per epoch, and plot each loss on its own axis.\n" ] }, { "cell_type": "code", "execution_count": null, "id": "b6ae0a2f", "metadata": {}, "outputs": [], "source": [ "Y_fm_onehot = np.eye(10)[y_fm_tr]\n", "rng = np.random.default_rng(SEED)\n", "fm_square = ... ### TODO_STUDENT ### the same 784-256-10 net, but with sigmoid outputs and the square loss\n", "hist_fm_sq = train_scratch(fm_square, X_fm_tr, y_fm_tr, lr=0.1, epochs=10, batch_size=256, rng=rng, desc=\"Fashion-MNIST, square loss\",\n", " X_test=X_fm_te, y_test=y_fm_te, targets=Y_fm_onehot)\n", "\n", "fig, axes = plt.subplots(1, 3, figsize=(14, 3))\n", "axes[0].plot(range(1, 11), hist_fm_ce[\"test_acc\"], \"o-\", label=\"softmax + cross-entropy\")\n", "axes[0].plot(range(1, 11), hist_fm_sq[\"test_acc\"], \"s-\", label=\"sigmoid + square loss\")\n", "axes[0].set_xlabel(\"epoch\"); axes[0].set_ylabel(\"test accuracy\"); axes[0].legend()\n", "for ax, h, name in [(axes[1], hist_fm_ce, \"cross-entropy\"), (axes[2], hist_fm_sq, \"square loss\")]:\n", " ax.plot(np.convolve(h[\"loss\"], np.ones(50) / 50, mode=\"valid\"))\n", " ax.set_xlabel(\"iteration\"); ax.set_ylabel(name + \" (moving avg)\")\n", "plt.show()\n", "print(\"test accuracy per epoch\")\n", "print(\" cross-entropy:\", np.round(hist_fm_ce[\"test_acc\"], 3))\n", "print(\" square loss: \", np.round(hist_fm_sq[\"test_acc\"], 3))\n", "\n", "# The output-layer gradient for ONE confidently wrong output: true class, but score o = -5\n", "o = -5.0\n", "p_softmax_like = 1 / (1 + np.exp(-o)) # the probability the model gives the true class (~0.007)\n", "print(f\"\\nconfidently wrong output (o = {o}):\")\n", "print(f\" cross-entropy gradient p - 1 = {p_softmax_like - 1:+.4f}\")\n", "print(f\" square-loss gradient (s - 1) * s * (1 - s) = {(p_softmax_like - 1) * p_softmax_like * (1 - p_softmax_like):+.4f}\")\n" ] }, { "cell_type": "markdown", "id": "d0ffa0c2", "metadata": {}, "source": [ "**Written answer (B).** In 2–3 sentences, explain the difference using the output-layer gradient: compare the softmax/cross-entropy gradient from (A) with what the sigmoid's derivative does to the square-loss gradient when an output is confidently wrong.\n", "\n", "*How to think about it.* Write the gradient that reaches one output unit's score in each case: (p − t) for softmax + cross-entropy, and (z − t)·z(1 − z) for sigmoid + square loss. Plug in a confidently wrong output, e.g. target t = 1 but output 0.01. How large is each gradient? Early in training many outputs are wrong: which loss gives those units a strong push? (The by-hand example in the NN_BASICS note, §7, shows the same effect with numbers.)\n", "\n", "*Your answer here.*\n" ] }, { "cell_type": "markdown", "id": "627b106b", "metadata": {}, "source": [ "### (C) Using PyTorch / d2l [15 points]\n", "\n", "The same 784-256-10 architecture as a d2l `Classifier` (d2l supplies the cross-entropy loss, the\n", "accuracy, and plain SGD with learning rate `lr`), trained by `d2l.Trainer` with the same settings as\n", "(A). PyTorch's default `nn.Linear` initialization is uniform in ±1/√n_in, the same as our `Linear`, so\n", "the two networks start from statistically identical weights (though not the same random draws).\n" ] }, { "cell_type": "code", "execution_count": null, "id": "35a2747a", "metadata": {}, "outputs": [], "source": [ "class ArrayData(d2l.DataModule):\n", " \"\"\"In-memory tensors as a d2l DataModule: train = shuffled minibatches, val = the test set.\"\"\"\n", "\n", " def __init__(self, X_train, y_train, X_val, y_val, batch_size):\n", " super().__init__()\n", " self.save_hyperparameters()\n", "\n", " def get_dataloader(self, train):\n", " tensors = (self.X_train, self.y_train) if train else (self.X_val, self.y_val)\n", " return self.get_tensorloader(tensors, train)\n", "\n", "\n", "class MLPClassifier(d2l.Classifier):\n", " \"\"\"Part I's network from nn layers: Flatten -> Linear -> ReLU -> Linear (softmax is inside the loss).\"\"\"\n", "\n", " def __init__(self, n_in, n_hidden, n_out, lr):\n", " super().__init__()\n", " self.save_hyperparameters()\n", " self.net = ... ### TODO_STUDENT ### the same architecture as (A), from nn layers\n", "\n", "\n", "def accuracy_torch(model, X, y):\n", " model.eval()\n", " with torch.no_grad():\n", " return (model(X).argmax(dim=1) == y).float().mean().item()\n", "\n", "\n", "# --- Fashion-MNIST ---\n", "torch.manual_seed(SEED)\n", "fm_torch = MLPClassifier(784, 256, 10, lr=0.1)\n", "fm_data = ArrayData(Xt_fm_tr, yt_fm_tr, Xt_fm_te, yt_fm_te, batch_size=256)\n", "t0 = time.time()\n", "... ### TODO_STUDENT ### train for 10 epochs with d2l's Trainer (num_gpus=0: CPU)\n", "fm_torch_time = time.time() - t0\n", "\n", "# --- Wine (13-16-3; the progress plot is switched off for this tiny run) ---\n", "torch.manual_seed(SEED)\n", "wine_torch = MLPClassifier(13, 16, 3, lr=0.1)\n", "wine_torch.board.display = False\n", "wine_data = ArrayData(torch.tensor(X_wine_tr, dtype=torch.float32), torch.tensor(y_wine_tr),\n", " torch.tensor(X_wine_te, dtype=torch.float32), torch.tensor(y_wine_te), batch_size=16)\n", "d2l.Trainer(max_epochs=100, num_gpus=0).fit(wine_torch, wine_data)\n" ] }, { "cell_type": "code", "execution_count": null, "id": "519a7ae3", "metadata": {}, "outputs": [], "source": [ "# d2l's live plot switches matplotlib to SVG output; switch back to PNG (SVG scatter plots of\n", "# thousands of points make the notebook huge)\n", "from matplotlib_inline.backend_inline import set_matplotlib_formats\n", "set_matplotlib_formats(\"png\")\n", "\n", "Xw_tr, Xw_te = torch.tensor(X_wine_tr, dtype=torch.float32), torch.tensor(X_wine_te, dtype=torch.float32)\n", "rows = [\n", " (\"Wine (k=16)\", \"from scratch\", *wine_scratch[16]),\n", " (\"Wine (k=16)\", \"PyTorch/d2l\", accuracy_torch(wine_torch, Xw_tr, torch.tensor(y_wine_tr)),\n", " accuracy_torch(wine_torch, Xw_te, torch.tensor(y_wine_te))),\n", " (\"Fashion-MNIST\", \"from scratch\", hist_fm_ce[\"train_acc\"][-1], hist_fm_ce[\"test_acc\"][-1]),\n", " (\"Fashion-MNIST\", \"PyTorch/d2l\", accuracy_torch(fm_torch, Xt_fm_tr, yt_fm_tr), accuracy_torch(fm_torch, Xt_fm_te, yt_fm_te)),\n", "]\n", "print(f\"{'dataset':15s} {'implementation':15s} {'train':>7s} {'test':>7s}\")\n", "for r in rows:\n", " print(f\"{r[0]:15s} {r[1]:15s} {r[2]:7.3f} {r[3]:7.3f}\")\n", "print(f\"\\n(d2l Trainer, Fashion-MNIST, 10 epochs on CPU: {fm_torch_time:.1f} s including the live plot)\")\n", "gap = abs(rows[2][3] - rows[3][3])\n", "print(f\"from-scratch vs PyTorch test-accuracy difference on Fashion-MNIST: {gap:.3f}\", \"(OK: within 2%)\" if gap <= 0.02 else \"(> 2%: look for a bug)\")\n" ] }, { "cell_type": "markdown", "id": "a1f5018b", "metadata": {}, "source": [ "### Part II — Train it well\n", "\n", "Everything from here on is in PyTorch. The helper below is the plain PyTorch training loop — the\n", "same loop `d2l.Trainer` runs for you — returning the training loss at every iteration and train/test\n", "(and optionally validation) accuracy after every epoch.\n" ] }, { "cell_type": "code", "execution_count": null, "id": "13026025", "metadata": {}, "outputs": [], "source": [ "def train_and_log(net, optimizer, X, y, X_test, y_test, epochs, batch_size=256, X_val=None, y_val=None, seed=SEED,\n", " desc=\"train\"):\n", " loss_fn = nn.CrossEntropyLoss()\n", " g = torch.Generator().manual_seed(seed)\n", " hist = {\"loss\": [], \"train_acc\": [], \"test_acc\": [], \"val_acc\": []}\n", " bar = progress(epochs * -(-len(X) // batch_size), desc)\n", " for epoch in range(epochs):\n", " net.train() # dropout ON\n", " perm = torch.randperm(len(X), generator=g)\n", " for start in range(0, len(X), batch_size):\n", " idx = perm[start:start + batch_size]\n", " loss = loss_fn(net(X[idx]), y[idx])\n", " optimizer.zero_grad()\n", " loss.backward()\n", " optimizer.step()\n", " hist[\"loss\"].append(loss.item())\n", " bar.update(1)\n", " hist[\"train_acc\"].append(accuracy_torch(net, X, y)) # accuracy_torch switches to eval mode (dropout OFF)\n", " hist[\"test_acc\"].append(accuracy_torch(net, X_test, y_test))\n", " bar.set_postfix(epoch=epoch + 1, test_acc=f\"{hist['test_acc'][-1]:.3f}\")\n", " if X_val is not None:\n", " hist[\"val_acc\"].append(accuracy_torch(net, X_val, y_val))\n", " bar.close()\n", " return hist\n", "\n", "\n", "def mlp(n_in=784, n_hidden=256, n_out=10):\n", " return nn.Sequential(nn.Linear(n_in, n_hidden), nn.ReLU(), nn.Linear(n_hidden, n_out))\n" ] }, { "cell_type": "markdown", "id": "0bb0a3c1", "metadata": {}, "source": [ "### (D) Optimizers and the learning rate [15 points]\n", "\n", "Five runs of the same network, 5 epochs each: plain SGD at lr = 0.01, 0.1, 1.0; SGD with momentum\n", "0.9 (lr = 0.1); Adam (lr = 10⁻³). Problem 6 builds on these curves.\n" ] }, { "cell_type": "code", "execution_count": null, "id": "d2f5cc4f", "metadata": {}, "outputs": [], "source": [ "OPTIMIZERS = [\"SGD lr=0.01\", \"SGD lr=0.1\", \"SGD lr=1.0\", \"SGD+momentum 0.9, lr=0.1\", \"Adam lr=1e-3\"]\n", "\n", "\n", "def make_optimizer(name, params):\n", " if name == \"SGD lr=0.01\":\n", " return ... ### TODO_STUDENT ### plain SGD, lr = 0.01\n", " if name == \"SGD lr=0.1\":\n", " return ... ### TODO_STUDENT ### plain SGD, lr = 0.1\n", " if name == \"SGD lr=1.0\":\n", " return ... ### TODO_STUDENT ### plain SGD, lr = 1.0\n", " if name == \"SGD+momentum 0.9, lr=0.1\":\n", " return ... ### TODO_STUDENT ### SGD with momentum 0.9, lr = 0.1\n", " if name == \"Adam lr=1e-3\":\n", " return ... ### TODO_STUDENT ### Adam, lr = 1e-3\n", " raise ValueError(name)\n", "\n", "\n", "opt_hist = {}\n", "for name in OPTIMIZERS:\n", " torch.manual_seed(SEED)\n", " net = mlp()\n", " opt_hist[name] = train_and_log(net, make_optimizer(name, net.parameters()), Xt_fm_tr, yt_fm_tr,\n", " Xt_fm_te, yt_fm_te, epochs=5, desc=name)\n", "\n", "fig, axes = plt.subplots(1, 2, figsize=(13, 3.8))\n", "for name in OPTIMIZERS:\n", " L = np.array(opt_hist[name][\"loss\"])\n", " axes[0].plot(np.convolve(L, np.ones(20) / 20, mode=\"valid\"), label=name, lw=1.2)\n", " axes[1].plot(L[:300], label=name, lw=0.8)\n", "axes[0].set_yscale(\"log\"); axes[0].set_xlabel(\"iteration\"); axes[0].set_ylabel(\"training loss (moving avg of 20, log)\")\n", "axes[0].set_title(\"all 5 epochs\"); axes[0].legend(fontsize=8)\n", "axes[1].set_xlabel(\"iteration\"); axes[1].set_ylabel(\"training loss\"); axes[1].set_title(\"first 300 iterations, raw\")\n", "plt.show()\n", "print(f\"{'optimizer':28s} test accuracy after 5 epochs\")\n", "for name in OPTIMIZERS:\n", " print(f\"{name:28s} {opt_hist[name]['test_acc'][-1]:.3f}\")\n" ] }, { "cell_type": "markdown", "id": "14c879f2", "metadata": {}, "source": [ "### (E) Regularization: weight decay and dropout [20 points]\n", "\n", "To make overfitting visible: a wider network (784-1024-1024-10) on a fixed **5,000-image** subset,\n", "trained with Adam (lr = 10⁻³) for 40 epochs, long enough that training accuracy clearly exceeds test\n", "accuracy. A separate 5,000-image **validation** set (also from the training images) is used to pick\n", "the weight-decay strength, so the test set is never used for a choice.\n" ] }, { "cell_type": "code", "execution_count": null, "id": "09dae72d", "metadata": {}, "outputs": [], "source": [ "def dropout_layer(X, p):\n", " \"\"\"Inverted dropout, training mode: zero each entry with probability p, scale the survivors by 1/(1-p).\"\"\"\n", " assert 0 <= p <= 1\n", " if p == 1:\n", " return torch.zeros_like(X)\n", " mask = ... ### TODO_STUDENT ### keep each entry independently with probability 1 - p\n", " return ... ### TODO_STUDENT ### zero the dropped entries and rescale the survivors\n", "\n", "\n", "class DropoutScratch(nn.Module):\n", " def __init__(self, p):\n", " super().__init__()\n", " self.p = p\n", "\n", " def forward(self, X):\n", " return ... ### TODO_STUDENT ### dropout in training mode only; identity in evaluation mode\n", "\n", "\n", "# CHECKPOINT: statistics vs nn.Dropout, and identity in eval mode\n", "torch.manual_seed(SEED)\n", "X_ones = torch.ones(2000, 500)\n", "for p in [0.2, 0.5, 0.8]:\n", " mine, ref = DropoutScratch(p).train()(X_ones), nn.Dropout(p).train()(X_ones)\n", " zeros_mine, zeros_ref = (mine == 0).float().mean().item(), (ref == 0).float().mean().item()\n", " print(f\" p={p}: fraction zeroed yours {zeros_mine:.3f} / nn.Dropout {zeros_ref:.3f}; mean yours {mine.mean().item():.3f} / nn.Dropout {ref.mean().item():.3f}\")\n", " assert abs(zeros_mine - p) < 0.01, \"the fraction of zeroed entries should be about p\"\n", " assert abs(mine.mean().item() - 1.0) < 0.02, \"the mean should be preserved (did you rescale by 1/(1-p)?)\"\n", "assert torch.equal(DropoutScratch(0.5).eval()(X_ones), X_ones), \"in eval mode dropout must return its input unchanged\"\n", "print(\"CHECKPOINT PASSED\")\n" ] }, { "cell_type": "code", "execution_count": null, "id": "03d2c43b", "metadata": {}, "outputs": [], "source": [ "g = torch.Generator().manual_seed(1)\n", "perm = torch.randperm(len(Xt_fm_tr), generator=g)\n", "X_small, y_small = Xt_fm_tr[perm[:5000]], yt_fm_tr[perm[:5000]]\n", "X_val, y_val = Xt_fm_tr[perm[5000:10000]], yt_fm_tr[perm[5000:10000]]\n", "\n", "\n", "def wide_mlp(dropout=None):\n", " layers = [nn.Linear(784, 1024), nn.ReLU()]\n", " if dropout:\n", " layers.append(DropoutScratch(dropout))\n", " layers += [nn.Linear(1024, 1024), nn.ReLU()]\n", " if dropout:\n", " layers.append(DropoutScratch(dropout))\n", " layers.append(nn.Linear(1024, 10))\n", " return nn.Sequential(*layers)\n", "\n", "\n", "def run_regularized(dropout=None, weight_decay=0.0, epochs=40):\n", " torch.manual_seed(SEED)\n", " net = wide_mlp(dropout)\n", " opt = ... ### TODO_STUDENT ### Adam, lr=1e-3, with the L2 penalty passed as weight_decay\n", " desc = f\"dropout p={dropout}\" if dropout else (f\"weight decay {weight_decay:g}\" if weight_decay else \"no regularization\")\n", " return train_and_log(net, opt, X_small, y_small, Xt_fm_te, yt_fm_te, epochs=epochs, batch_size=128,\n", " X_val=X_val, y_val=y_val, desc=desc)\n", "\n", "\n", "t0 = time.time()\n", "reg_hist = {\"no regularization\": run_regularized()}\n", "for wd in [1e-4, 1e-3, 3e-3]:\n", " reg_hist[f\"weight decay {wd:g}\"] = run_regularized(weight_decay=wd)\n", "best_wd = max([k for k in reg_hist if k.startswith(\"weight\")], key=lambda k: reg_hist[k][\"val_acc\"][-1])\n", "reg_hist[\"dropout p=0.5\"] = run_regularized(dropout=0.5)\n", "print(f\"5 runs x 40 epochs: {time.time() - t0:.0f} s; best weight decay by VALIDATION accuracy: {best_wd}\")\n", "\n", "print(f\"\\n{'run':22s} {'train':>6s} {'val':>6s} {'test':>6s} {'gap (train-test)':>17s}\")\n", "for k, h in reg_hist.items():\n", " print(f\"{k:22s} {h['train_acc'][-1]:6.3f} {h['val_acc'][-1]:6.3f} {h['test_acc'][-1]:6.3f} {h['train_acc'][-1] - h['test_acc'][-1]:17.3f}\")\n", "\n", "plt.figure(figsize=(7, 4))\n", "for k, color in [(\"no regularization\", \"C0\"), (best_wd, \"C1\"), (\"dropout p=0.5\", \"C2\")]:\n", " ep = range(1, len(reg_hist[k][\"train_acc\"]) + 1)\n", " plt.plot(ep, reg_hist[k][\"train_acc\"], \"-\", color=color, label=f\"{k}: train\")\n", " plt.plot(ep, reg_hist[k][\"test_acc\"], \"--\", color=color, label=f\"{k}: test\")\n", "plt.xlabel(\"epoch\"); plt.ylabel(\"accuracy\"); plt.legend(fontsize=8); plt.title(\"5,000 training images, 784-1024-1024-10\")\n", "plt.show()\n" ] }, { "cell_type": "markdown", "id": "b1927790", "metadata": {}, "source": [ "**Reading the table (E).** The table reports the final train/test gap for each run. Note what you see (which runs shrink the gap, and which improve test accuracy); Problem 5 asks you to explain it.\n", "\n", "*How to think about it.* Read one column at a time. The gap is train minus test: which runs make it smaller? Now the test column alone: which runs actually raise it? A gap can shrink because test accuracy went *up* or because train accuracy went *down*. Which happened for each regularizer?\n", "\n", "*Your answer here.*\n" ] }, { "cell_type": "markdown", "id": "c78e5788", "metadata": {}, "source": [ "### (F) Initialization and vanishing gradients [15 points]\n", "\n", "A 10-hidden-layer MLP (width 128) in four variants: {sigmoid, ReLU} × {small fixed-scale init\n", "N(0, 0.01²), the init matched to the activation (Xavier for sigmoid, He/Kaiming for ReLU)}. First the\n", "gradient norm of every layer's weights after one backward pass on one batch; then 2 epochs of\n", "training each.\n" ] }, { "cell_type": "code", "execution_count": null, "id": "0c4d52da", "metadata": {}, "outputs": [], "source": [ "def deep_mlp(activation, init, depth=10, width=128):\n", " layers, n_in = [], 784\n", " for _ in range(depth):\n", " lin = nn.Linear(n_in, width)\n", " n_in = width\n", " if init == \"small\":\n", " ... ### TODO_STUDENT ### small fixed-scale init, N(0, 0.01^2)\n", " elif activation == \"sigmoid\":\n", " ... ### TODO_STUDENT ### Xavier/Glorot init (matched to sigmoid)\n", " else:\n", " ... ### TODO_STUDENT ### He/Kaiming init (matched to ReLU)\n", " nn.init.zeros_(lin.bias)\n", " layers += [lin, nn.Sigmoid() if activation == \"sigmoid\" else nn.ReLU()]\n", " layers.append(nn.Linear(width, 10))\n", " return nn.Sequential(*layers)\n", "\n", "\n", "def layer_grad_norms(net, X, y):\n", " net.zero_grad()\n", " loss = ... ### TODO_STUDENT ### one forward pass and the loss on this batch\n", " ... ### TODO_STUDENT ### one backward pass\n", " return [m.weight.grad.norm().item() for m in net if isinstance(m, nn.Linear)]\n", "\n", "\n", "VARIANTS = [(\"sigmoid\", \"small\"), (\"sigmoid\", \"scaled\"), (\"relu\", \"small\"), (\"relu\", \"scaled\")]\n", "LABEL = {(\"sigmoid\", \"small\"): \"sigmoid, N(0, 0.01^2)\", (\"sigmoid\", \"scaled\"): \"sigmoid, Xavier\",\n", " (\"relu\", \"small\"): \"ReLU, N(0, 0.01^2)\", (\"relu\", \"scaled\"): \"ReLU, He/Kaiming\"}\n", "plt.figure(figsize=(7, 4))\n", "for v in VARIANTS:\n", " torch.manual_seed(SEED)\n", " norms = layer_grad_norms(deep_mlp(*v), Xt_fm_tr[:256], yt_fm_tr[:256])\n", " plt.semilogy(range(1, len(norms) + 1), norms, \"o-\", label=LABEL[v])\n", "plt.xlabel(\"layer (1 = first, next to the input; 11 = output layer)\"); plt.ylabel(\"||dL/dW|| (log scale)\")\n", "plt.title(\"gradient norm per layer, one batch, before training\"); plt.legend(fontsize=8); plt.show()\n", "\n", "print(\"2 epochs of SGD (lr=0.1) each:\")\n", "for v in VARIANTS:\n", " torch.manual_seed(SEED)\n", " net = deep_mlp(*v)\n", " h = train_and_log(net, torch.optim.SGD(net.parameters(), lr=0.1), Xt_fm_tr, yt_fm_tr, Xt_fm_te, yt_fm_te, epochs=2,\n", " desc=LABEL[v])\n", " print(f\" {LABEL[v]:24s} test accuracy {h['test_acc'][-1]:.3f} (10 classes: chance = 0.100)\")\n" ] }, { "cell_type": "markdown", "id": "eacb786d", "metadata": {}, "source": [ "**Written answer (F).** Which variants learn at all? In 2–3 sentences, connect the gradient-norm plot to the chain rule you implemented in Problem 1: what happens to a product of many per-layer factors that are each smaller than 1?\n", "\n", "*How to think about it.* In Problem 1, each sigmoid layer's backward pass multiplied the incoming gradient by σ'(net). What is the largest value σ' can ever take? Multiply ten such numbers. For ReLU, what is the derivative on the units that are active? Then think about the weights: every layer's backward pass also multiplies by W, so what does a tiny initial scale such as 0.01 do to the signal after ten layers, in the forward pass and in the backward pass? Which of your four curves stays roughly level across layers?\n", "\n", "*Your answer here.*\n" ] }, { "cell_type": "markdown", "id": "9c52f2d1", "metadata": {}, "source": [ "---\n", "## Problem 3 — Variational Autoencoder (VAE), using PyTorch [50 points]\n", "\n", "Problem 1's autoencoder compresses and reconstructs, but nothing organizes its codes and there's no\n", "principled way to generate a new point. A VAE makes the latent space a probability distribution\n", "shaped like the prior N(0, I). Data: MNIST, 784 pixel values in [0, 1]. See the VAE lecture note.\n", "\n", "### (A) Build and train it [20 points]\n", "\n", "The architecture is given: a shared encoder trunk with two linear heads (μ and log σ²), and a\n", "decoder 2 → 400 → 784 with a sigmoid output. Read it; the two-heads-on-one-trunk shape is the only\n", "new piece. You implement what makes it a VAE: the reparameterization, the KL term, and the loss.\n" ] }, { "cell_type": "code", "execution_count": null, "id": "4f7f55fa", "metadata": {}, "outputs": [], "source": [ "class VAE(nn.Module):\n", " def __init__(self, n_in=784, n_hidden=400, n_latent=2):\n", " super().__init__()\n", " self.trunk = nn.Sequential(nn.Linear(n_in, n_hidden), nn.ReLU()) # shared encoder trunk\n", " self.mu_head = nn.Linear(n_hidden, n_latent) # head 1: the mean mu(x)\n", " self.logvar_head = nn.Linear(n_hidden, n_latent) # head 2: log sigma^2(x)\n", " self.decoder = nn.Sequential(nn.Linear(n_latent, n_hidden), nn.ReLU(), # 2 -> hidden -> 784,\n", " nn.Linear(n_hidden, n_in), nn.Sigmoid()) # pixel intensities in [0, 1]\n", "\n", " def encode(self, x):\n", " h = self.trunk(x)\n", " return self.mu_head(h), self.logvar_head(h)\n", "\n", " def reparameterize(self, mu, logvar):\n", " eps = ... ### TODO_STUDENT ### eps ~ N(0, I), same shape as mu\n", " return ... ### TODO_STUDENT ### z = mu + sigma * eps, with sigma = exp(logvar / 2)\n", "\n", " def forward(self, x):\n", " mu, logvar = self.encode(x)\n", " z = self.reparameterize(mu, logvar)\n", " return self.decoder(z), mu, logvar\n", "\n", "\n", "def kl_to_standard_normal(mu, logvar):\n", " \"\"\"KL( N(mu, diag(sigma^2)) || N(0, I) ): one value per row.\"\"\"\n", " return ... ### TODO_STUDENT ### the closed-form Gaussian KL, summed over the latent dimensions\n", "\n", "\n", "# CHECKPOINT: closed-form KL vs a Monte-Carlo estimate E_q[log q(z) - log p(z)]\n", "torch.manual_seed(SEED)\n", "mu_chk, logvar_chk = torch.randn(500, 2), 0.5 * torch.randn(500, 2)\n", "q = torch.distributions.Normal(mu_chk, torch.exp(0.5 * logvar_chk))\n", "p = torch.distributions.Normal(0.0, 1.0)\n", "z = q.sample((20000,))\n", "kl_mc = (q.log_prob(z) - p.log_prob(z)).sum(-1).mean(0)\n", "kl_closed = kl_to_standard_normal(mu_chk, logvar_chk)\n", "rel = (abs(kl_closed.sum() - kl_mc.sum()) / kl_closed.sum()).item()\n", "print(f\"total KL over 500 random (mu, logvar): closed form {kl_closed.sum().item():.1f}, Monte Carlo {kl_mc.sum().item():.1f}, relative difference {rel:.4f}\")\n", "assert kl_closed.shape == (500,), \"return one KL value per row\"\n", "assert rel < 0.01, \"closed-form KL and the Monte-Carlo estimate should agree to about 1%\"\n", "print(\"CHECKPOINT PASSED\")\n" ] }, { "cell_type": "markdown", "id": "0a281487", "metadata": {}, "source": [ "**Written answer (A)(ii): why sampling z directly would block the gradient.**\n", "\n", "*How to think about it.* Backprop needs, at every step between the loss and the encoder's weights, an answer to \"if I nudge this input a little, how does this output change?\" Ask that question about the step \"draw z at random from N(μ, σ²)\": if you nudge μ, what happens to one particular random draw? Is there a derivative to pass back? Now ask it about z = μ + σ·ε, where ε was drawn first and is held fixed: what is ∂z/∂μ, and ∂z/∂σ? Where did the randomness go?\n", "\n", "*Your answer here.*\n" ] }, { "cell_type": "code", "execution_count": null, "id": "bba7cb5c", "metadata": {}, "outputs": [], "source": [ "def vae_loss(x, x_hat, mu, logvar):\n", " recon = ... ### TODO_STUDENT ### binary cross-entropy summed over pixels, averaged over the batch\n", " kl = ... ### TODO_STUDENT ### the KL term, averaged over the batch\n", " return recon, kl\n", "\n", "\n", "def train_vae(model, X, epochs=10, batch_size=128, lr=1e-3, use_kl=True, seed=SEED, desc=\"VAE\"):\n", " torch.manual_seed(seed)\n", " opt = torch.optim.Adam(model.parameters(), lr=lr)\n", " hist = {\"recon\": [], \"kl\": []}\n", " bar = progress(epochs * -(-len(X) // batch_size), desc)\n", " for epoch in range(epochs):\n", " model.train()\n", " perm = torch.randperm(len(X))\n", " tot_r = tot_k = 0.0\n", " for start in range(0, len(X), batch_size):\n", " xb = X[perm[start:start + batch_size]]\n", " x_hat, mu, logvar = model(xb)\n", " recon, kl = vae_loss(xb, x_hat, mu, logvar)\n", " loss = recon + kl if use_kl else recon\n", " opt.zero_grad()\n", " loss.backward()\n", " opt.step()\n", " tot_r += recon.item() * len(xb)\n", " tot_k += kl.item() * len(xb)\n", " bar.update(1)\n", " hist[\"recon\"].append(tot_r / len(X))\n", " hist[\"kl\"].append(tot_k / len(X))\n", " bar.set_postfix(epoch=epoch + 1, recon=f\"{hist['recon'][-1]:.1f}\", kl=f\"{hist['kl'][-1]:.1f}\")\n", " bar.close()\n", " return hist\n", "\n", "\n", "torch.manual_seed(SEED)\n", "vae = VAE()\n", "t0 = time.time()\n", "vae_hist = train_vae(vae, Xt_mn_tr, epochs=10)\n", "print(f\"VAE, 10 epochs on CPU: {time.time() - t0:.0f} s; final reconstruction {vae_hist['recon'][-1]:.1f}, KL {vae_hist['kl'][-1]:.2f} (nats per image)\")\n", "\n", "fig, axes = plt.subplots(1, 2, figsize=(9, 3))\n", "axes[0].plot(range(1, 11), vae_hist[\"recon\"], \"o-\"); axes[0].set_title(\"reconstruction (BCE)\"); axes[0].set_xlabel(\"epoch\")\n", "axes[1].plot(range(1, 11), vae_hist[\"kl\"], \"o-\", color=\"C1\"); axes[1].set_title(\"KL( q(z|x) || N(0,I) )\"); axes[1].set_xlabel(\"epoch\")\n", "plt.show()\n" ] }, { "cell_type": "markdown", "id": "89cee280", "metadata": {}, "source": [ "### (B) Explore the latent space [15 points]\n", "\n", "The comparison model is given: a *plain* autoencoder with exactly your VAE's architecture, but with\n", "the sampling and the KL term removed (its code for x is simply μ(x)).\n" ] }, { "cell_type": "code", "execution_count": null, "id": "6c9ea334", "metadata": {}, "outputs": [], "source": [ "class PlainAE(VAE):\n", " \"\"\"Your VAE with the sampling and the KL term removed: code(x) = mu(x), trained on reconstruction only.\"\"\"\n", "\n", " def forward(self, x):\n", " mu, logvar = self.encode(x)\n", " return self.decoder(mu), mu, logvar\n", "\n", "\n", "torch.manual_seed(SEED)\n", "plain_ae = PlainAE()\n", "train_vae(plain_ae, Xt_mn_tr, epochs=10, use_kl=False, desc=\"plain AE\")\n", "\n", "# (i) organization: the 2-D codes of the 10,000 test digits\n", "with torch.no_grad():\n", " mu_vae = vae.encode(Xt_mn_te)[0].numpy()\n", " mu_ae = plain_ae.encode(Xt_mn_te)[0].numpy()\n", "fig, axes = plt.subplots(1, 2, figsize=(12, 5))\n", "for ax, codes, title in [(axes[0], mu_vae, \"VAE: latent means mu(x)\"), (axes[1], mu_ae, \"plain autoencoder: codes\")]:\n", " sc = ax.scatter(codes[:, 0], codes[:, 1], c=y_mn_te, cmap=\"tab10\", s=2, rasterized=True)\n", " ax.set_title(title); ax.set_aspect(\"equal\")\n", "fig.colorbar(sc, ax=axes, ticks=range(10), label=\"digit\")\n", "plt.show()\n", "print(\"spread of the codes (std per dimension): VAE\", mu_vae.std(0).round(2), \" plain AE\", mu_ae.std(0).round(2))\n" ] }, { "cell_type": "code", "execution_count": null, "id": "c3ab81d2", "metadata": {}, "outputs": [], "source": [ "def show_grid(images, nrow, title, size=1.0):\n", " ncol = int(np.ceil(len(images) / nrow))\n", " fig, axes = plt.subplots(nrow, ncol, figsize=(ncol * size, nrow * size))\n", " for ax, img in zip(np.ravel(axes), images):\n", " ax.imshow(img.reshape(28, 28), cmap=\"gray\")\n", " for ax in np.ravel(axes):\n", " ax.axis(\"off\")\n", " fig.suptitle(title)\n", " plt.show()\n", "\n", "\n", "vae.eval()\n", "with torch.no_grad():\n", " # (ii) generation: z straight from the prior, no input image\n", " torch.manual_seed(SEED)\n", " z = ... ### TODO_STUDENT ### 64 draws from the prior N(0, I)\n", " samples = vae.decoder(z).numpy()\n", " # the latent \"map\": decode a regular 15 x 15 grid over [-2, 2]^2\n", " ticks = np.linspace(-2, 2, 15)\n", " zz = torch.tensor([[a, b] for b in ticks[::-1] for a in ticks], dtype=torch.float32)\n", " latent_map = vae.decoder(zz).numpy()\n", "show_grid(samples, 8, \"64 VAE samples from z ~ N(0, I)\", size=0.8)\n", "show_grid(latent_map, 15, \"decoded 15 x 15 grid of z over [-2, 2]^2\", size=0.55)\n", "\n", "# (iii) interpolation between two test digits of different classes\n", "i_a, i_b = int(np.where(y_mn_te == 1)[0][0]), int(np.where(y_mn_te == 0)[0][0])\n", "with torch.no_grad():\n", " z_a, z_b = vae.encode(Xt_mn_te[[i_a]])[0], vae.encode(Xt_mn_te[[i_b]])[0]\n", " steps = []\n", " for t in np.linspace(0, 1, 10):\n", " z_line = ... ### TODO_STUDENT ### linear interpolation between the two latent means\n", " steps.append(vae.decoder(z_line).numpy()[0])\n", "show_grid(steps, 1, f\"interpolating from a {y_mn_te[i_a]} to a {y_mn_te[i_b]} in latent space\", size=1.1)\n" ] }, { "cell_type": "markdown", "id": "37cd40e0", "metadata": {}, "source": [ "**Written answer (B).** Comment briefly: how do the VAE's latent codes differ from the plain autoencoder's, and what do the samples, the latent map, and the interpolation show?\n", "\n", "*How to think about it.* Compare the two scatter plots' *axes* first: the range of the codes, and where the digits sit relative to (0, 0). Then imagine drawing z from N(0, I), a cloud of radius about 2 around the origin: in each plot, would those draws land on digit codes or in empty space? Which term of the VAE's loss pulls the codes toward that cloud? For the latent map and the interpolation, look at the regions *between* digit clusters: does the decoder still produce something digit-like there, and why would it, given how the VAE was trained?\n", "\n", "*Your answer here.*\n" ] }, { "cell_type": "markdown", "id": "75ce8a3f", "metadata": {}, "source": [ "### (C) Written only [15 points]\n" ] }, { "cell_type": "markdown", "id": "84c69b71", "metadata": {}, "source": [ "**Written answer (C).** Why can EM's E-step be computed exactly for a Gaussian mixture but not for this VAE's decoder, and why is a learned encoder the right fix (rather than, say, just running more EM iterations)?\n", "\n", "*How to think about it.* An E-step computes the posterior p(z | x) ∝ p(x | z) p(z) and normalizes it. For a Gaussian mixture, z takes one of K values: how many numbers do you compute, and how do you normalize them? For the VAE, z is a continuous 2-D vector and p(x | z) is a neural network: what would normalizing require (think of a sum over K values turning into an integral over all z)? Is running more iterations a fix for a step you cannot carry out even once? Finally, count the work: 60,000 training images, each needing its own posterior, at every iteration and again for every new image, versus one network that outputs a guess for any x in one forward pass. The VAE note's remark \"Why not just run EM longer?\" is about exactly this.\n", "\n", "*Your answer here.*\n" ] }, { "cell_type": "markdown", "id": "f7209195", "metadata": {}, "source": [ "---\n", "## Problem 4 — Generative Adversarial Network (GAN): Run It and Explain It [10 points]\n", "\n", "HW4 fit a Gaussian mixture to `2gaussian.txt` by maximizing a likelihood; Problem 3 maximizes a\n", "lower bound on one; a GAN never writes down a density. **All the code in this problem is given**:\n", "both networks (small MLPs with LeakyReLU), their optimizers (Adam with β₁ = 0.5, the usual GAN\n", "setting), the alternating training step, and the plots. Run it, read the training step until you\n", "can explain every line, and answer the questions from what you see. The data are standardized\n", "(zero mean, unit variance per coordinate).\n" ] }, { "cell_type": "code", "execution_count": null, "id": "2bb8e5e1", "metadata": {}, "outputs": [], "source": [ "def standardize(A):\n", " return (A - A.mean(axis=0)) / A.std(axis=0)\n", "\n", "\n", "real2 = torch.tensor(standardize(gauss2), dtype=torch.float32)\n", "# the true component means (from HW4 Problem 3), in the same standardized coordinates\n", "means2 = (np.array([[3.0, 3.0], [7.0, 4.0]]) - gauss2.mean(axis=0)) / gauss2.std(axis=0)\n", "Z_DIM = 2\n", "\n", "\n", "def make_gan(hidden=64, lr=2e-4, seed=SEED):\n", " torch.manual_seed(seed)\n", " G = nn.Sequential(nn.Linear(Z_DIM, hidden), nn.LeakyReLU(0.2), nn.Linear(hidden, hidden), nn.LeakyReLU(0.2),\n", " nn.Linear(hidden, 2)) # noise -> 2-D point (raw values, no squashing)\n", " D = nn.Sequential(nn.Linear(2, hidden), nn.LeakyReLU(0.2), nn.Linear(hidden, hidden), nn.LeakyReLU(0.2),\n", " nn.Linear(hidden, 1)) # 2-D point -> raw score (the loss applies the sigmoid)\n", " opt_G = torch.optim.Adam(G.parameters(), lr=lr, betas=(0.5, 0.999))\n", " opt_D = torch.optim.Adam(D.parameters(), lr=lr, betas=(0.5, 0.999))\n", " return G, D, opt_G, opt_D\n", "\n", "\n", "bce = nn.BCEWithLogitsLoss()\n" ] }, { "cell_type": "markdown", "id": "942f5261", "metadata": {}, "source": [ "### (A) Two modes: how a GAN trains [5 points]\n", "\n", "The training step below alternates one D update and one G update. Read both functions carefully:\n", "question (A)(i) asks about the two lines marked ◀.\n" ] }, { "cell_type": "code", "execution_count": null, "id": "9ebe3bf0", "metadata": {}, "outputs": [], "source": [ "def d_step(G, D, opt_D, real):\n", " n = len(real)\n", " fake = G(torch.randn(n, Z_DIM)).detach() # ◀ generated points, detached\n", " loss_D = bce(D(real), torch.ones(n, 1)) + bce(D(fake), torch.zeros(n, 1)) # real -> label 1, fake -> label 0\n", " opt_D.zero_grad()\n", " loss_D.backward()\n", " opt_D.step()\n", " return loss_D.item()\n", "\n", "\n", "def g_step(G, D, opt_G, n):\n", " fake = G(torch.randn(n, Z_DIM)) # fresh points, NOT detached\n", " loss_G = bce(D(fake), torch.ones(n, 1)) # ◀ label 1: \"call these real\"\n", " opt_G.zero_grad()\n", " loss_G.backward()\n", " opt_G.step()\n", " return loss_G.item()\n", "\n", "\n", "def train_gan(real, steps=4000, batch_size=256, d_steps=1, g_steps=1, checkpoints=(0, 500, 2000, 4000), desc=\"GAN\",\n", " **gan_kwargs):\n", " G, D, opt_G, opt_D = make_gan(**gan_kwargs)\n", " snapshots, hist = {}, {\"D\": [], \"G\": []}\n", " bar = progress(steps, desc, unit=\"step\")\n", " for step in range(steps + 1):\n", " if step in checkpoints:\n", " with torch.no_grad():\n", " snapshots[step] = G(torch.randn(3000, Z_DIM)).numpy()\n", " if step == steps:\n", " break\n", " for _ in range(d_steps):\n", " loss_D = d_step(G, D, opt_D, real[torch.randint(len(real), (batch_size,))])\n", " for _ in range(g_steps):\n", " loss_G = g_step(G, D, opt_G, batch_size)\n", " hist[\"D\"].append(loss_D)\n", " hist[\"G\"].append(loss_G)\n", " bar.update(1)\n", " bar.close()\n", " return G, snapshots, hist\n", "\n", "\n", "def mode_fractions(samples, means):\n", " \"\"\"Fraction of generated points closest to each true component mean.\"\"\"\n", " d = ((samples[:, None, :] - means[None, :, :]) ** 2).sum(-1)\n", " return np.bincount(d.argmin(axis=1), minlength=len(means)) / len(samples)\n", "\n", "\n", "def mode_spreads(samples, means):\n", " \"\"\"Per-mode standard deviation (x, y) of the points closest to each mean -- checks each mode's SHAPE,\n", " not just its share: a mode collapsed onto a thin line has the right count but a tiny spread.\"\"\"\n", " lab = ((samples[:, None, :] - means[None, :, :]) ** 2).sum(-1).argmin(axis=1)\n", " return [samples[lab == k].std(axis=0) if (lab == k).sum() > 1 else np.zeros(2) for k in range(len(means))]\n", "\n", "\n", "t0 = time.time()\n", "G2, snaps2, gan_hist = train_gan(real2)\n", "print(f\"GAN, 4000 steps: {time.time() - t0:.1f} s\")\n" ] }, { "cell_type": "code", "execution_count": null, "id": "854893aa", "metadata": {}, "outputs": [], "source": [ "fig, axes = plt.subplots(1, len(snaps2), figsize=(4 * len(snaps2), 3.6), sharex=True, sharey=True)\n", "for ax, (step, s) in zip(axes, snaps2.items()):\n", " ax.scatter(real2[:, 0], real2[:, 1], s=2, c=\"lightgray\", label=\"real\")\n", " ax.scatter(s[:, 0], s[:, 1], s=2, c=\"C3\", alpha=0.5, label=\"generated\")\n", " ax.set_title(f\"step {step}\")\n", "axes[0].legend(markerscale=5)\n", "plt.show()\n", "\n", "plt.figure(figsize=(8, 3))\n", "w = 50\n", "plt.plot(np.convolve(gan_hist[\"D\"], np.ones(w) / w, mode=\"valid\"), label=\"D loss (real + fake terms)\")\n", "plt.plot(np.convolve(gan_hist[\"G\"], np.ones(w) / w, mode=\"valid\"), label=\"G loss\")\n", "plt.axhline(2 * np.log(2), ls=\"--\", c=\"k\", lw=0.8, label=\"2 ln 2 = 1.386 (D can't tell)\")\n", "plt.axhline(np.log(2), ls=\":\", c=\"k\", lw=0.8, label=\"ln 2 = 0.693\")\n", "plt.xlabel(\"step\"); plt.ylabel(f\"loss (moving avg of {w})\"); plt.legend(fontsize=8); plt.show()\n", "\n", "print(\"fraction of generated points nearest each true mode:\", mode_fractions(snaps2[4000], means2).round(3),\n", " \" (data: 2000 / 4000 points -> 0.333 / 0.667)\")\n", "real2_np = real2.numpy()\n", "for k, (sg, sr) in enumerate(zip(mode_spreads(snaps2[4000], means2), mode_spreads(real2_np, means2))):\n", " print(f\"mode {k}: spread (std x, y) generated {sg.round(2)} vs real {sr.round(2)}\")\n" ] }, { "cell_type": "markdown", "id": "d66bec37", "metadata": {}, "source": [ "**Written answer (A).** (i) Why is `fake` detached in `d_step` but not in `g_step`, and why does `g_step` use label 1? (ii) Where does D's loss settle, and why is that the value when D can't tell real from fake? (iii) What does \"implicit density\" mean here, and why can you still sample from G?\n", "\n", "*How to think about it.* (i) Ask what each step is *supposed* to change. In the D step, whose weights should move? If gradients from D's loss did flow into G, which way would they push G: toward fooling D, or toward being caught? For the label, compare the two G losses early in training, when D easily spots fakes (D(G(z)) ≈ 0.01): what is the slope of log(1 − D) there, and what is the slope of −log D? Which one gives G a useful push when G is doing badly? (ii) If D outputs 1/2 for every point, what is −log(1/2)? D's loss is the sum of two such terms, one for real and one for fake points. Compare with the dashed lines on your plot. (iii) To *draw* a sample from G, what do you need to do? To *evaluate* p(x) at some given point x, what would you need to know about G: which z's map to x, and how G stretches or squeezes space there?\n", "\n", "*Your answer here.*\n" ] }, { "cell_type": "markdown", "id": "974a7199", "metadata": {}, "source": [ "### (B) Three modes: when a GAN fails [5 points]\n", "\n", "The same GAN on `3gaussian.txt` (components of 2000, 3000, 5000 points), under three schedules: the\n", "default (1 D step per G step), G 5× stronger, and D 5× stronger. For each run the cell prints the\n", "fraction of samples nearest each true mode and each mode's spread, next to the real data's. A mode\n", "can have the right *share* of points but the wrong *shape*: check both.\n" ] }, { "cell_type": "code", "execution_count": null, "id": "dcf31b67", "metadata": {}, "outputs": [], "source": [ "real3 = torch.tensor(standardize(gauss3), dtype=torch.float32)\n", "means3 = (np.array([[3.0, 3.0], [7.0, 4.0], [5.0, 7.0]]) - gauss3.mean(axis=0)) / gauss3.std(axis=0)\n", "\n", "settings = {\"default (1 D step : 1 G step)\": dict(),\n", " \"G overpowers D (1 D : 5 G steps)\": dict(g_steps=5),\n", " \"D stronger (5 D : 1 G steps)\": dict(d_steps=5)}\n", "fig, axes = plt.subplots(1, len(settings), figsize=(4.2 * len(settings), 3.8), sharex=True, sharey=True)\n", "for ax, (name, kw) in zip(axes, settings.items()):\n", " t0 = time.time()\n", " _, snaps, _ = train_gan(real3, checkpoints=(4000,), desc=name.split(\" (\")[0], **kw)\n", " frac = mode_fractions(snaps[4000], means3)\n", " ax.scatter(real3[:, 0], real3[:, 1], s=2, c=\"lightgray\")\n", " ax.scatter(snaps[4000][:, 0], snaps[4000][:, 1], s=2, c=\"C3\", alpha=0.5)\n", " ax.set_title(f\"{name}\\nmode fractions {frac.round(2)}\", fontsize=9)\n", " print(f\"{name:34s} fractions nearest each mode {frac.round(3)} (data 0.2 / 0.3 / 0.5) {time.time() - t0:.0f} s\")\n", " print(f\"{'':34s} per-mode spread (std x, y): \" + \" \".join(f\"({a:.2f}, {b:.2f})\" for a, b in mode_spreads(snaps[4000], means3)))\n", "plt.show()\n", "print(f\"{'real data':34s} per-mode spread (std x, y): \" + \" \".join(f\"({a:.2f}, {b:.2f})\" for a, b in mode_spreads(real3.numpy(), means3)))\n" ] }, { "cell_type": "markdown", "id": "fd958210", "metadata": {}, "source": [ "**Written answer (B).** What happens under each schedule (use the printed shares *and* spreads)? Explain mode collapse in terms of what G is rewarded for. Why is HW4's EM far less prone to this failure?\n", "\n", "*How to think about it.* Use the printout. The shares tell you how much of G's output lands near each mode (compare with 0.2 / 0.3 / 0.5); the spreads tell you how *wide* each generated mode is (compare with the real data's row). A run can have the right shares and still squeeze a mode too narrow. For the mechanism, look at G's loss: it asks only \"does the *current* D think my points are real?\" Is there any term that penalizes G for leaving a mode empty? What happens if G gets 5 moves before D can react, and what changes when D gets 5 moves first? For EM, every real point contributes log p(x) to its objective: what happens to that sum if a whole cluster has no component near it? (See the GAN note's \"Common Failure Modes\", especially \"Who moves faster\".)\n", "\n", "*Your answer here.*\n" ] }, { "cell_type": "markdown", "id": "e3e950ea", "metadata": {}, "source": [ "---\n", "## Problem 5 (THEORY) — Why Dropout Lowers Train Accuracy but Raises Test Accuracy [25 points]\n", "\n", "In Problem 2(E), adding dropout most likely *lowered* your training accuracy, raised (or held) your\n", "test accuracy, and shrank the train/test gap. (If your run didn't show this clearly, say so and\n", "explain why your setup might not show it.) In a short written answer, explain why dropout produces\n", "this pattern. Then explain (conceptually) why dropout is turned *off* at test time rather than left\n", "on, and what the 1/(1−p) scaling in your inverted-dropout implementation is for.\n" ] }, { "cell_type": "markdown", "id": "507496b2", "metadata": {}, "source": [ "**Answer.**\n", "\n", "*How to think about it.* During training, a network with dropout can never count on any particular hidden unit being present. What does that force each unit to learn, compared with a network whose units can rely on a few partners always being there? Why would that hurt on the training set but help on new data? For test time: would you rather predict with one random thinned network, or with the full network? And if each unit was present only a fraction (1 − p) of the time during training, what happens to the size of the next layer's input when all units are suddenly present, unless something is rescaled?\n", "\n", "*Your answer here.*\n" ] }, { "cell_type": "markdown", "id": "1d84a470", "metadata": {}, "source": [ "---\n", "## Problem 6 (THEORY) — Diagnosing the Learning Rate from a Loss Curve [25 points]\n", "\n", "Two students train the same network with gradient descent and plot training loss vs. iteration.\n", "Student A's curve decreases for a few iterations, then starts oscillating wildly and trending upward.\n", "Student B's curve decreases smoothly but extremely slowly, flattening out after many iterations at a\n", "loss still far above what a properly-tuned run achieves. In a short written answer, diagnose each\n", "student's problem in terms of the learning rate and explain the *mechanism* behind each pattern (not\n", "just \"it's too big/small\"; think of the step size relative to the curvature of the loss). Point to\n", "which of your own Problem 2(D) curves, if any, look like A's or B's.\n", "\n", "A third student, C, trains with minibatch SGD. C's curve drops quickly, then settles into a band that\n", "keeps jittering up and down, and it gets no narrower or lower even when C trains ten times longer.\n", "Nothing diverges, and C's loss evaluated once per epoch on the *whole* training set is smooth. Where\n", "does the jitter come from? (It is not the curvature story above.) Name two changes that would let C's\n", "loss settle lower, and what each one costs. Which of your raw Problem 2(D) curves show a band like this?\n" ] }, { "cell_type": "markdown", "id": "64342fd2", "metadata": {}, "source": [ "**Answer.**\n", "\n", "*How to think about it.* Use the 1-D picture from the NN_BASICS note (§13.1): one gradient step multiplies the distance to the minimum by (1 − ηλ), where λ is the curvature. What happens to that factor, and to the distance, when ηλ > 2? When ηλ is tiny? For Student C: a minibatch gradient is an *estimate* built from a few hundred examples. Near the minimum, where the true gradient is almost zero, what is left in the estimate? With a fixed step size, can that leftover ever die out, and what would you have to change for it to?\n", "\n", "*Your answer here.*\n" ] }, { "cell_type": "markdown", "id": "b627e906", "metadata": {}, "source": [ "---\n", "## OPTIONAL PROBLEMS [no credit]\n" ] }, { "cell_type": "markdown", "id": "bb2074ba", "metadata": {}, "source": [ "## Problem 8 — Backprop Equations in Matrix Form [optional, no credit]\n", "\n", "For Problem 2's network on a minibatch X (n rows): Z₁ = XW₁ + b₁, H = ReLU(Z₁), O = HW₂ + b₂,\n", "P = softmax(O), and G = (P − Y)/n with Y the one-hot labels. Then\n", "\n", "* dW₂ = Hᵀ G, db₂ = 1ᵀ G (column sums of G)\n", "* dH = G W₂ᵀ, dZ₁ = dH ⊙ 1[Z₁ > 0]\n", "* dW₁ = Xᵀ dZ₁, db₁ = 1ᵀ dZ₁\n", "\n", "The cell below evaluates these four expressions directly and checks them against your layer\n", "classes.\n" ] }, { "cell_type": "code", "execution_count": null, "id": "813d8cae", "metadata": {}, "outputs": [], "source": [ "rng = np.random.default_rng(SEED)\n", "net9 = Network([Linear(13, 8, rng), ReLU(), Linear(8, 3, rng)], SoftmaxCrossEntropy())\n", "Xb, yb = X_wine_tr[:32], y_wine_tr[:32]\n", "net9.loss.forward(net9.forward(Xb), yb)\n", "net9.backward()\n", "L1, L2 = net9.linear_layers()\n", "n = len(Xb)\n", "Z1 = Xb @ L1.W + L1.b\n", "Hh = np.maximum(Z1, 0)\n", "P = net9.loss.p\n", "Y = np.eye(3)[yb]\n", "Gm = ... ### TODO_STUDENT ### (optional) dL/dO in matrix form\n", "dW2 = ... ### TODO_STUDENT ### (optional) dL/dW2\n", "db2 = ... ### TODO_STUDENT ### (optional) dL/db2\n", "dZ1 = ... ### TODO_STUDENT ### (optional) dL/dZ1 = (G W2^T) masked by ReLU'\n", "dW1 = ... ### TODO_STUDENT ### (optional) dL/dW1\n", "db1 = ... ### TODO_STUDENT ### (optional) dL/db1\n", "for name, mine, layer_grad in [(\"dW2\", dW2, L2.dW), (\"db2\", db2, L2.db), (\"dW1\", dW1, L1.dW), (\"db1\", db1, L1.db)]:\n", " assert np.allclose(mine, layer_grad), name\n", " print(f\"{name}: matrix form matches the layer classes (max abs diff {np.abs(mine - layer_grad).max():.1e})\")\n" ] } ], "metadata": { "kernelspec": { "display_name": "Python 3", "language": "python", "name": "python3" } }, "nbformat": 4, "nbformat_minor": 5 }