A guy on Reddit said I could fine-tune a 7B model on a single 3090, and he was dead wrong
Back in March I picked up a used 3090 for $550 because some guy on r/LocalLLaMA swore up and down that 24GB of VRAM was plenty for fine-tuning a 7B model. I loaded up my dataset, about 12,000 rows of support tickets from my job, and hit run. Two hours later I got an out of memory error, then tried every quantized config I could find and still couldn't get past a few hundred steps. Turns out the optimizer states alone eat up most of the card unless you do some serious tricks like LoRA with tiny rank, and even then my batches were so small the training was garbage. I burned a full weekend and ended up renting an A100 on a cloud service for $1.89 an hour just to finish the job in one evening. If you're thinking about getting into local fine-tuning because of a forum comment, please check the actual VRAM math first. What's the smallest setup you've personally gotten to work for fine-tuning anything bigger than a 3B model?