My AI side project kept crashing on Sunday nights and I finally found the tiny bug
I built a little text classifier bot in October that labels posts for a subreddit, and it kept dying every Sunday around 2am for 3 weeks straight. Turns out my cron job and my weekly model retrain were both hitting the same Hugging Face endpoint at the same time, so the API rate limit was killing the bot mid batch. I added a 10 minute delay and a simple retry loop with backoff, and it has run clean for 12 days now. Honestly I felt dumb because I spent $6 on extra GPU time before checking the logs. Anyone else run into rate limit stuff when you schedule retraining and inference on the same key?