The technology of voice and sound. From text-to-speech systems and voice cloning to speech recognition and audio engineering, this channel covers the cutting edge of how machines learn to speak, listen, and sound convincingly human.
#5573: Audio Post-Processing for AI Podcast Pipelines
Silence trimming, loudness normalization, EQ, and tempo fixes — the CLI tools that turn raw TTS output into a finished episode.
#5570: Why Phone Trees Break and AI Voice Agents Can Fix Them
AI is good at routing calls and bad at resolving them. That single distinction explains why menu-replacing voice agents work and human-replacing on...
#5551: Why Learning Makes Us Happier at Any Age
The credential isn't what makes learning feel good — choosing it is. What the research says about self-directed learning, aging, and hard times.
#5545: Jet Engines on Bicycles: The Physics of Bad Ideas
A pulsejet strapped to a bike sounds like a fighter jet — and burns 55 gallons an hour to move one guy who could have pedaled.
#5487: Why Your Spreadsheet Mangles José's Name
A deep dive into character sets, from ASCII to UTF-8, and why José becomes "José" in your spreadsheet.
#5485: Fine-Tuning Parakeet for Hebrew and Your Own Jargon
NVIDIA's Parakeet beats Whisper on Android — but can you teach it Hebrew, or just your own jargon? Two answers, one much happier.
#5466: Hebrew Words Hidden in English Text
Daniel wants a classifier that spots Hebrew written in Latin letters — and it turns out nobody's built one.
#5465: Packaging AI Pipelines So They Actually Get Reused
Daniel's podcast pipeline works — so why can't he reuse it? Recipes, containers, and Hebrew code-switching TTS.
#5464: Your Keyboard's Hidden Data Problem
Your keyboard knows your email address, your phrases, your habits — and you can't take any of it with you.
#5461: TTS Can't Pronounce Hebrew Inside English
Your TTS reads Hebrew words with English phonetics. Here's why — and why the obvious fix doesn't work yet.
#5460: Four Small Models, One Android Phone: Does It Actually Work?
A chained on-device dictation pipeline — VAD, ASR, cleanup — and why "it feels smooth" isn't the same as knowing it works.
#5459: Rebuilding the Podcast: Turn-Taking, Buttons, and Chatterbox
Daniel wants a push-to-interrupt button and a live voice loop. The research says the button is easy and the live part is a research project.
#5456: Parakeet vs Whisper: Picking a Phone ASR Model
Why Whisper loses on Android, why Parakeet v2 beat v3, and how to benchmark speech-to-text without any tooling.
#5435: When Your TTS Model Eats the Numbers
Numbers, dates, and acronyms break text-to-speech in specific, documented ways. Here's where normalization lives — and why it depends on your model.
#5430: Two Boxes: ASR and the Text Fixer Behind It
Punctuation, casing, ITN, disfluency — the four-job layer between raw ASR output and text you can actually read.
#5422: Switching Android Keyboards Without the Tap Dance
One listener wants a one-tap jump between his Parakeet voice keyboard and his typing keyboard. Android says no — unless you know the trick.
#5413: Who Actually Picks Your In-Flight Movie?
Eight in ten passengers use seatback screens — but at Southwest, one person picks everything. Inside the tiny teams behind airline entertainment.
#5397: Chaining Small Models for Voice Cleanup
Six cleanup stages at 97% accuracy each compound to 83% end-to-end. So how many small models can you actually chain?
#5396: Teaching a Small Model to Stop Spelling Out Numbers
Your ASR pipeline is fine until someone dictates "three point two" and gets "three point two" spelled out. Here's how inverse text normalization ac...
#5390: Chaining Small Models for Dictation Cleanup
Daniel's Android dictation fork won't render "three point five" as a decimal. How many models does cleanup actually need?