Setup for the Neural-Network Homeworks (HW5–HW7)

HW5–HW7 use PyTorch plus the helper package of the Dive into Deep Learning (d2l) book. You set it up once, before HW5, in about 20 minutes. You do not need a GPU. Every homework runs on a normal laptop; a GPU just makes HW6–HW7 faster.

1. Install

In a terminal (macOS/Linux):

python3 -m venv ~/cs6140_env
source ~/cs6140_env/bin/activate        # repeat this line in every new terminal
pip install --upgrade pip
pip install torch torchvision numpy scipy pandas matplotlib scikit-learn jupyterlab requests tqdm sacrebleu gensim
pip install d2l==1.0.3 --no-deps        # keep --no-deps (see below)
jupyter lab

Windows (Command Prompt): the same, except create and activate the environment with py -m venv %USERPROFILE%\cs6140_env and %USERPROFILE%\cs6140_env\Scripts\activate. If you have an NVIDIA GPU, see "NVIDIA GPU" below before installing torch.

Why --no-deps? The d2l package pins very old versions of numpy/matplotlib that won't install on current Python. d2l itself works fine with current versions, so we install it without its pins.

2. Check that it works

Paste into a notebook cell and run:

import torch
from d2l import torch as d2l
device = 'cuda' if torch.cuda.is_available() else 'mps' if torch.backends.mps.is_available() else 'cpu'
print('torch', torch.__version__, '| device:', device)
data = d2l.FashionMNIST(batch_size=256)          # downloads ~30 MB the first time
print(len(data.train), 'training images')

Expected: no errors, a device name (cuda = NVIDIA GPU, mps = Apple Silicon GPU, cpu = no GPU, which is fine), and 60000 training images. That's it: you're set for HW5–HW7. Each starter notebook begins with a cell that picks your device automatically.

3. Which machine do I have?

Your machineWhat happens
Mac with Apple Silicon (M1–M4)Uses the Mac's GPU (mps) automatically. Nothing extra to install.
Linux or Windows with an NVIDIA GPUUses the GPU (cuda). On Windows this needs a special torch install (below).
Anything else (Intel Mac, laptop without NVIDIA GPU)Runs on the CPU. HW5 is fast; for HW6–HW7 use the starters' FAST setting, or Colab (below).

Optional details

Open these only if they apply to you.

NVIDIA GPU (Windows, Linux, WSL2)
Apple Silicon notes
Google Colab (free GPU in the browser)
Using plain PyTorch instead of d2l's Trainer

The starters use d2l's Module / Trainer classes, which match the book. You may use a plain PyTorch training loop instead (the most common style in practice), as long as you evaluate exactly as the starter does. This is what d2l.Trainer does for you:

model = model.to(device)
opt = torch.optim.SGD(model.parameters(), lr=0.1)
loss_fn = torch.nn.CrossEntropyLoss()
for epoch in range(num_epochs):
    model.train()                                  # dropout / batch norm in training mode
    for X, y in train_loader:
        X, y = X.to(device), y.to(device)
        loss = loss_fn(model(X), y)
        opt.zero_grad(); loss.backward(); opt.step()
    model.eval()                                   # evaluation mode
    with torch.no_grad():
        val_acc = sum((model(X.to(device)).argmax(1) == y.to(device)).float().mean().item()
                      for X, y in val_loader) / len(val_loader)
    print(f'epoch {epoch}: val acc {val_acc:.3f}')

No other frameworks (Lightning, Keras, JAX, Hugging Face) are needed for this course.

conda users, and picking the right Jupyter kernel
Troubleshooting
SymptomFix
Installing d2l tries to build old numpy and fails You left out --no-deps. Run pip install --upgrade numpy matplotlib pandas scipy, then the d2l line again.
ModuleNotFoundError: torch in Jupyter Wrong kernel; see the section above.
Out of memory, or the machine freezes during training Halve batch_size, use the FAST setting, and restart the kernel to free old models.
DataLoader worker ... exited unexpectedly (usually in .py scripts) Put the script's code under if __name__ == "__main__":, or use num_workers=0.
OMP: Error #15 ... libomp already initialized (conda on Mac) Use the plain venv from Step 1 instead of conda.
Numbers differ slightly between runs or machines Normal (random initialization, floating point). Compare with tolerances.
What each homework downloads

d2l stores downloads in a ../data folder next to the notebook's folder. Keep your HW notebooks in sibling folders (e.g. cs6140/hw5, cs6140/hw6) so they share one download cache.

Still stuck? Post on Piazza with your OS, the output of the Step 2 check, and the full error message (as text, not a screenshot).