PyTorch
The deep learning framework that won research, and increasingly production too.
What it is
PyTorch is an open-source deep learning framework for Python that provides dynamic computation graphs, GPU acceleration, and a flexible platform for building neural networks and machine learning models.
PyTorch provides tensors, autograd for automatic differentiation, neural network modules (torch.nn), optimizers, and utilities for loading and preprocessing data. It supports CPU and GPU computations seamlessly.
- Best known for
- Define-by-run graphs that behave like ordinary Python and debug like it
- Licence
- BSD 3-clause
- Maintained by
- The PyTorch Foundation under the Linux Foundation
When to use it
The question documentation cannot answer for you — because it cannot recommend something else.
Reach for it when
- Training or fine-tuning neural networks, especially where you want to inspect and modify the model
- Research and experimentation — the dynamic graph makes debugging straightforward
- Working with the Hugging Face ecosystem, which is PyTorch-first
Look elsewhere when
- The problem is tabular and a gradient-boosted tree would do it better with less effort
- You need a mature mobile or embedded deployment path — TensorFlow Lite is more established there
Installation
pip install torch torchvision torchaudioGetting started
The smallest useful thing you can do with it, and what each part means.
import torch
x = torch.tensor([[1,2],[3,4]])
y = torch.rand(2,2)
print(x + y)import torch
x = torch.rand(2,3)
y = torch.rand(3,2)
print(torch.mm(x, y))Advanced usage
Where the library earns its place over a simpler alternative.
import torch.nn as nn
import torch.nn.functional as F
class Net(nn.Module):
def __init__(self):
super(Net, self).__init__()
self.fc1 = nn.Linear(10, 5)
self.fc2 = nn.Linear(5, 1)
def forward(self, x):
x = F.relu(self.fc1(x))
x = torch.sigmoid(self.fc2(x))
return x
model = Net()import torch.optim as optim
x = torch.rand(100,10)
y = torch.rand(100,1)
criterion = nn.MSELoss()
optimizer = optim.SGD(model.parameters(), lr=0.01)
for epoch in range(10):
optimizer.zero_grad()
outputs = model(x)
loss = criterion(outputs, y)
loss.backward()
optimizer.step()
print(f'Epoch {epoch+1}, Loss: {loss.item()}')device = torch.device('cuda' if torch.cuda.is_available() else 'cpu')
model.to(device)
x, y = x.to(device), y.to(device)
outputs = model(x)x = torch.tensor([1.0,2.0,3.0], requires_grad=True)
y = x.pow(2).sum()
y.backward()
print(x.grad)Errors and fixes
The failures you are most likely to hit, and what actually resolves them.
- RuntimeError: CUDA out of memory
- Reduce batch size or move computations to CPU if GPU memory is insufficient.
- RuntimeError: shape mismatch
- Ensure the input and target tensors have compatible shapes for the model and loss function.
- ModuleNotFoundError: No module named 'torch'
- Install PyTorch using pip or conda in the current Python environment.
Best practices
- Use `torch.nn.Module` to structure neural networks cleanly.
- Leverage `torch.utils.data.DataLoader` for batching and shuffling datasets.
- Always zero gradients with `optimizer.zero_grad()` before backpropagation.
- Move tensors to GPU using `.to(device)` for faster training.
- Use PyTorch Lightning or similar frameworks for cleaner training loops in production.
Alternatives
Comparable options, and the reason you would pick one over the other.
Background
Why it exists, and what it was reacting to.
PyTorch was developed by Facebook's AI Research lab (FAIR) and released in 2016. It was designed to provide a more intuitive and flexible framework than TensorFlow at the time, supporting dynamic computation graphs that make debugging and experimentation easier. PyTorch quickly became popular in both research and production environments.
