llama.cpp
large language model inferencing library written in C++
Episodes
-
#5673: Why Lowering Temperature Broke Daniel's ScriptsLowering temperature should make output safer, not incoherent. Daniel's DeepSeek scripts say otherwise — and the reason may be stranger than he thi... -
#5485: Fine-Tuning Parakeet for Hebrew and Your Own JargonNVIDIA's Parakeet beats Whisper on Android — but can you teach it Hebrew, or just your own jargon? Two answers, one much happier. -
#5436: Small Models as Rewriters, Not WritersWhy "don't say X" prompts backfire, and how a tiny grammar-constrained model can scrub a script without breaking its grammar. -
#5384: Android ASR Runtimes: LiteRT, ExecuTorch, and Why Your Phone Has No VRAMWhy does your phone have no VRAM number? A tour of Android's runtime layer and what it takes to run ASR locally. -
#5407: Hemmingway-1 and the War on WaffleA 27B model promises answers without the preamble. Its benchmark is homegrown — and the behavior it targets has a paper trail. -
#5393: Fine-Tuning a Model on 100 Hand-Edited AnswersYou don't need 10,000 examples to make a model sound like you. The real number is closer to 100 — if the edits are opinionated.