Direct Preference Optimization
method for training language models directly on preference data without a separate reward model; replaces reinforcement learning from human feedback with a classification-style objective on preferred vs dispreferred outputs
Episodes
-
#5517: What 5,000 AI-Written Podcast Episodes RevealDaniel archived every AI-generated episode with the model that wrote it. Now he wants to know which models loop — and how to catch it. -
#5412: Editing vs. Note-Taking for AI Fine-TunesHand-editing a model's output gives three training signals at once. Writing notes gives one — and a weaker one at that.