Stop letting your GPU idle while your CPU struggles to feed it
Handling the Bottleneck with Prefetching and Parallelism
The most common mistake is loading data synchronously. When the model finishes a batch, the GPU sits idle while the CPU fetches the next chunk from the disk. You can kill this latency by using multi-process loading. In PyTorch, this is handled via the num_workers parameter in the DataLoader.
Setting num_workers to the number of CPU cores usually helps, but be careful with memory overhead. If you're using a massive dataset, you should combine this with pin_memory=True, which speeds up the transfer from CPU RAM to GPU VRAM by using page-locked memory.
Optimized Formats for Large Scale Training
Stop using raw CSVs or thousands of tiny JSON files. Opening and closing files creates massive overhead. For a real-world deployment, you need binary formats that support sequential reads and memory mapping.
- TFRecord: The gold standard for TensorFlow, storing data as a sequence of binary records.
- Apache Parquet: Incredible for tabular data due to columnar storage, which means you only load the features you actually need.
- WebDataset: Essential for vision tasks; it wraps data into POSIX tar files, allowing you to stream datasets over a network without needing to download the whole thing to a local SSD first.
A Practical Tutorial for Custom Data Pipelines
If you're building a custom LLM agent or a fine-tuning script, you'll likely need a custom dataset class. Here is a basic structure to ensure your data is preprocessed on the fly without blocking the training loop.
import torch
from torch.utils.data import Dataset, DataLoader
class EfficientDataset(Dataset):
def __init__(self, data_path):
# Load metadata or index files here, not the full dataset
self.data = self._load_index(data_path)
def __len__(self):
return len(self.data)
def __getitem__(self, idx):
# Perform heavy transformations here
sample = self.data[idx]
processed_sample = self.transform(sample)
return torch.tensor(processed_sample)
def transform(self, x):
# Example: Normalization or tokenization
return x / 255.0
# Deployment configuration for maximum throughput
loader = DataLoader(
dataset=EfficientDataset("data/train"),
batch_size=64,
shuffle=True,
num_workers=8,
pin_memory=True,
prefetch_factor=2
)Memory Mapping and Sharding
When your dataset exceeds your system RAM, memory mapping (mmap) is your best friend. It allows the OS to map a file directly into the virtual address space, loading pages only when they are accessed. For distributed training across multiple GPUs, you must implement sharding. This ensures that each GPU sees a unique subset of the data per epoch, preventing redundant computation and ensuring the gradient updates are based on a diverse sample of the global dataset. This is the only way to scale a deep dive project from a single local machine to a cluster.