A100
Episodes
-
#5411: Fine-Tuning at 4-Bit vs 16-Bit: What It Really CostsQLoRA cuts fine-tuning VRAM 15x and cost up to 85% — but you pay in training time, quality, and safety alignment. -
#5830: Serving Your Own Fine-Tuned Model in the CloudYou fine-tuned an open-weight model. Now how does anyone actually talk to it? Dedicated GPUs vs serverless inference, and the math that decides it. -
#5408: When Small NLP Models Beat the LLMFeature extraction, fill-mask, token classification — the classic NLP tasks still have a job. Here's when a small model beats a frontier API.