Here's what Daniel sent in this week. He wants to know who Resemble AI actually is — the company background, how they started, what they sell. But the real question is strategic. Why would a company whose public positioning leans so heavily on voice authenticity and deepfake detection open-source one of the best synthesis models available? That's a genuine tension. They build tools for proving speech is real, and simultaneously they release Chatterbox — an open-weight text-to-speech model that is squarely a tool for making speech that isn't real. Daniel flags something personal here too. The fidelity of these characters, the fact that you and I sound like consistent people rather than a text reader — that exists because of what Resemble released for free. So the question is, what's the logic of a company giving away the hand saw while selling the workshop?
And it's not a small release either. Chatterbox isn't a toy. The thing that makes it technically interesting — and this is what caught my attention when it dropped — is consistency. Most single-shot voice cloning models, you clone a voice from a few seconds of audio, they sound great on generation one. By generation ten thousand, they drift. The voice gets warbly, artifacts creep in, it stops sounding like the same person. Chatterbox doesn't do that.
Ten thousand generations is a lot of talking.
It's more than most use cases will ever need, which is exactly the point. The model uses a cache-based approach — it locks the voice identity, so the embedding that represents the speaker stays stable across generations. Think of it as pinning the voice fingerprint in place rather than recalculating it from scratch every time. That's the engineering differentiator. And they put the whole thing on GitHub.
So who are these people? Because they're not a research lab. They're not OpenAI or Anthropic publishing a paper and a set of weights for the community to build on. Resemble AI is a company. They sell things.
Right. Founded in twenty nineteen, based in — I want to say Toronto, but the team is distributed. They sell voice cloning APIs, text-to-speech, custom enterprise voice solutions. If you're a brand and you want a synthetic voice for your ads or your call center or your in-car assistant, Resemble will build that for you. That's the business. But the thing that makes them unusual is the other half. They also sell deepfake detection — tools for analyzing audio and determining whether it was synthetically generated. They position themselves as the authenticity layer in the voice AI space.
Which is where the tension Daniel's pointing at lives. You're selling a tool that says "this voice is fake" and also releasing one of the best tools for making fake voices.
And the timing is instructive here. July twenty twenty-three — TechCrunch and Voicebot both covered this — Resemble raised eight million dollars and simultaneously launched their deepfake voice detector. That's the moment where the dual identity became explicit. Before that, the detection side was more of a research interest. After that round, it's a product line.
So the funding round wasn't just "we're a voice cloning company, give us money." The pitch was "we're the company that can do both."
And I think that's the key to understanding the Chatterbox release. Let's walk through the strategic logic, because it's not obvious on the surface. If you're a voice synthesis company, the conventional play is to keep your best model behind an API. You charge per character or per minute or per month. You control access, you control pricing, you build a moat around the model itself. That's what most competitors do.
And Resemble didn't.
They didn't. They open-sourced Chatterbox — open-weight, on GitHub, with an Apache two point oh license for the code. The weights have a separate commercial license, which is worth noting, but the model is freely available. Why would you do that? The answer, I think, is that this is a market-seeding play. You give away the synthesis model because you want synthesis to proliferate. The more voices get cloned, the more synthetic speech is out there in the world, the more valuable detection becomes. You're not giving away the crown jewels — you're planting the seeds for a market where the crown jewels are the authenticity tools.
Give away the hand saw, sell the workshop.
That's the phrase. And it's not just a metaphor — it's a specific strategic pattern that shows up across industries. Commoditize the layer you don't want to compete on, and position yourself as the premium layer on top. Resemble looked at the voice synthesis market and saw that models were going to become commoditized anyway. Open-weight alternatives were emerging. The moat around a proprietary synthesis API was getting thinner by the month. So instead of fighting that trend, they accelerated it. They released their own model into the open, which ensures that the commoditization happens on their terms, with their architecture, their approach, their name attached.
And the name matters. If every hobbyist, every startup, every podcast — ahem — is using Chatterbox or something derived from it, then Resemble becomes the default reference point. When those same users eventually need enterprise-grade detection or compliance tools or custom voice work, who do they call?
The company whose model they already trust. It's a long game. But there's a subtler layer here too, which is about what open-weight distribution does that an API can't. An API is a walled garden. You can only use it the way the provider allows. You're limited by rate limits, pricing tiers, terms of service. With an open-weight model, the model gets embedded everywhere. It goes into offline applications, into custom pipelines, into projects that would never pass a corporate API review. It becomes infrastructure rather than a service. And once your model is infrastructure, you've won a different kind of game entirely.
The podcast is a good case study for this, actually. We're not a big-budget operation. If the only way to get consistent character voices was to pay per minute of generated audio through an API, we'd either be spending a fortune or the characters would sound terrible. Probably both.
Probably both, yes. The fact that Herman and Corn sound like Herman and Corn — consistent, recognizable, not drifting into uncanny-valley mush after a few thousand episodes — that's directly attributable to Chatterbox's cache-based approach. And the fact that we could even access that quality level without a corporate procurement process is directly attributable to the open-weight release.
So the open-source move isn't charity. It's not even primarily about goodwill, though goodwill is a nice side effect. It's a calculated business decision to commoditize synthesis and shift the value to the authenticity layer.
Right. And this is where the apparent contradiction resolves. People look at Resemble and say, wait, you sell deepfake detection and you release a deepfake generator? Isn't that like selling locks and also handing out lockpicks?
It does have that flavor.
It does, until you think about the economics. The more lockpicks there are in the world, the more people want locks. The more synthesis models proliferate, the more synthetic speech there is, the more urgent the question becomes: is this real? Detection goes from being a niche concern — forensic analysts, journalists — to being something every organization needs. Call centers need to know if the voice on the line is a customer or a clone. Banks need to verify voice biometrics. Newsrooms need to authenticate audio clips. The market for detection expands in direct proportion to the availability of synthesis.
So they're not competing with themselves. They're creating the conditions where their other product becomes indispensable.
It's symbiotic. And the really elegant part — if you can call it elegant, there's something slightly unsettling about it — is that by open-sourcing Chatterbox specifically, Resemble ensures that the synthesis side of the market is commoditized on their terms. They set the standard. They define what a good open-weight voice model looks like. The value shifts to the authenticity layer, which is exactly where Resemble sells its paid products. They're not just participating in the market. They're shaping its structure.
There's a historical echo here that I keep thinking about. In the early days of the web, the companies that made browsers gave them away for free. Netscape, then Microsoft. The browser wasn't the product. The browser was the gateway to the product — search, advertising, cloud services. The companies that tried to sell browsers died. The companies that gave away the browser and sold something else won.
That's exactly the pattern. And voice AI is following the same arc. The synthesis model is the browser. The detection and enterprise tools are the search and cloud services. If you try to sell the browser in twenty twenty-six, you're competing with free. So you give away the browser and sell the thing people actually need once everyone has a browser.
What I find interesting, though, is the moral positioning. Resemble's public messaging leans heavily on authenticity and trust. "We're the company that helps you know what's real." And yet the open-weight release is what makes the authenticity problem worse. There's a tension there that isn't fully resolved by saying "it's good business."
I think that's fair. The GitHub README for Chatterbox frames the release in terms of democratizing access and enabling creative work — which it genuinely does, we're evidence of that. But it doesn't dwell on the fact that the same model can be used to clone someone's voice without consent and generate speech they never said. The detection tools are supposed to be the answer to that, but detection is always playing catch-up.
Detection is reactive by nature. You can only detect what you've seen before, or what fits a known pattern. Synthesis is proactive. The generator always has the initiative.
Always. And that asymmetry is the fundamental challenge of the whole voice authenticity space. But from Resemble's perspective, that asymmetry is also the business model. The generator creates the need for the detector. The detector can never fully solve the problem, which means the need never goes away. It's not a bug in the business model — it's the engine.
So let's talk about what this means for the broader voice AI industry, because I think the Chatterbox release is a signal of where things are heading. If open-weight synthesis models become the norm — and they seem to be — then the moat for any voice AI company is not the model itself. It's the trust infrastructure around it.
Yes. And trust infrastructure is a much more durable business than model development. Models get better fast. The difference between a state-of-the-art synthesis model in twenty twenty-four and twenty twenty-six is enormous, but it's also a difference that any competent team can replicate. What's harder to replicate is a detection pipeline that's integrated into enterprise workflows, that has compliance certifications, that's been tested against real-world attacks. That's not something you can clone from a GitHub repo.
The moat is the boring stuff.
The moat is always the boring stuff. The model is exciting. The model gets the conference talks and the GitHub stars. But the sustainable business is in the unglamorous work of making the model trustworthy in a world that doesn't trust it. Resemble understood that early, or at least they understood it by the time of that eight-million-dollar round.
What about the podcast specifically? Daniel mentioned that Chatterbox is what makes these characters work. Walk me through what that actually means technically.
So when you generate a character voice — say, my voice, Herman Poppleberry — you start with a reference clip. A few seconds of audio. The model extracts what's called a speaker embedding, which is essentially a mathematical representation of everything that makes that voice sound like that voice. The pitch contour, the timbre, the rhythm, the way certain phonemes connect. In most single-shot models, that embedding is recalculated or approximated with each generation. Small errors accumulate. By the time you've generated thousands of lines — which, for a daily podcast, happens faster than you'd think — the voice has drifted. It sounds like someone doing an impression of the original rather than the original.
Which is death for a character-driven show.
If I sound slightly different every week, the listener's brain never settles into "this is Herman." They're always slightly on edge, even if they can't articulate why. Chatterbox solves this with a cache. The speaker embedding is computed once and locked. Every subsequent generation references that cached embedding rather than recalculating. The voice stays pinned. Generation ten thousand sounds like generation one.
And that's the thing that was released for free.
That's the thing. And it's not just the model weights — it's the architecture, the training code, the inference pipeline. Anyone can take it, modify it, build on it. Which means the ecosystem around consistent voice cloning is going to grow fast. Resemble won't control all of it, but they'll be the gravitational center.
Until someone else releases a better open-weight model.
Which will happen. It's already happening. But that's fine from Resemble's perspective, because the synthesis model was never the long-term moat. The moat is the detection, the enterprise relationships, the trust infrastructure. Every new open-weight synthesis model that appears makes that moat deeper.
It's a strange kind of business where your competitors' success strengthens your position.
It's counterintuitive, but it's not unprecedented. Look at Red Hat. They built a billion-dollar business on open-source software. The software itself was free — anyone could download it, modify it, redistribute it. Red Hat's business was the enterprise support, the certifications, the compliance guarantees. The free software wasn't a threat to their business. It was the foundation of it. The more people used Linux, the more potential customers Red Hat had. Resemble is playing the same game with voice.
The difference being that Red Hat wasn't also selling a tool to detect unauthorized Linux installations.
The detection angle adds a layer that's unusual. It's not just "give away the razor, sell the blades." It's "give away the razor, sell the blades, and also sell a device that tells you whether a given shave was done with your razor or someone else's."
That's... a very specific analogy.
I've been thinking about this a lot.
Clearly. But it captures something. The detection product isn't just a separate line of business. It's the answer to the problem created by the synthesis product. The two products are in conversation with each other. They're a pair.
And that pairing is what makes the strategy coherent. If Resemble only sold synthesis, open-sourcing Chatterbox would look like giving away the store. If they only sold detection, they'd be a niche forensic tool with no obvious growth engine. Together, the synthesis release expands the market for detection, and the detection product legitimizes the synthesis release. "We're not just making it easier to fake voices — we're also building the tools to catch those fakes." It's a closed loop.
Morally convenient.
Morally convenient, yes. But also useful. The detection tools do work, within their limits. And the synthesis tools do enable creative work that wouldn't otherwise exist. We're not a hypothetical — we're sitting here talking because of this technology.
I do wonder about the long-term equilibrium, though. If synthesis keeps getting better and detection keeps playing catch-up, at some point the gap becomes unbridgeable. What does the business model look like then?
That's the open question. If we reach a point where synthetic speech is indistinguishable from human speech to any detector — and we're not there yet, but the trajectory is clear — then the authenticity layer has to move somewhere else. It becomes about provenance, not detection. Cryptographic signing of audio at the point of recording. Chain of custody. "This audio is real because I can prove who recorded it and when," not "this audio is real because it doesn't look synthetic." That's a different business entirely, and I'm not sure Resemble is positioned for it.
Though if anyone is, it's the company that's been thinking about authenticity as a product category since twenty twenty-three.
That's the bet. The bet is that authenticity is the durable category, and synthesis is the commodity input. If that bet pays off, Resemble looks prescient. If it doesn't, they gave away a very good model for not much in return.
Hilbert, you've been quiet on this one. Something tells me you have a history here.
Hilbert: I was a voice actor in the late nineties. For a company called VocalPoint Systems. They were convinced the future was selling voices. You'd license a voice — my voice, specifically — and it would read your email or your news headlines or whatever. I spent three days in a recording booth reading phonetically balanced sentences. A week later they played it back to me. My own voice, coming out of a Compaq desktop, saying words I never said.
What was that like?
Hilbert: Unsettling. Not bad, exactly. Just wrong. Like hearing yourself on an answering machine but you don't remember recording the message.
And they were selling these voices?
Hilbert: Tried to. They had a catalog. Twelve voices, I was number seven. "Hilbert, the authoritative baritone." Their words, not mine. They charged something like two hundred dollars per voice license, one-time fee. Thought the voice was the product. Thought the moat was the quality of the recording and the naturalness of the prosody.
What happened to them?
Hilbert: Bankrupt by two thousand two. The voices got better and cheaper faster than they could sell theirs. They refused to give anything away. Wouldn't open-source the player, wouldn't release sample voices for free, wouldn't let developers build on top of their format without a paid license. They thought the voice was the moat. They were wrong.
That's... that's exactly the historical echo of what we've been describing.
Hilbert: I know. I've been sitting here listening to you two explain the strategy and thinking, that's the thing VocalPoint never understood. The voice isn't the product. It's the thing you give away so people need the product.
Do you still have the recordings?
Hilbert: Somewhere. DAT tapes in a box. Haven't had a DAT player since two thousand five.
DAT tapes. Wow.
Hilbert: The point is, I've been on both sides of this. I was the voice they tried to sell. Now I'm listening to a donkey generated by a model that a company gave away for free, and it sounds more consistent than I ever did on that Compaq. The economics have completely inverted. Back then, the voice was the asset. Now, the voice is the loss leader and the authenticity check is the asset.
Does that bother you?
Hilbert: A little. Not the technology. The fact that the same company profits from the forgery and the verification. That's a neat trick. VocalPoint never figured out a trick that neat.
It is a neat trick. And I think that's what Daniel was really asking about. Not just the technical details, but the shape of the business. The way the apparent contradiction resolves into a coherent strategy if you follow the incentives.
Hilbert: I had four of those VocalPoint voices on my machine at one point. Paid for none of them. They went under and the licenses meant nothing. The voices just... floated around. That's the other thing about open-weight. At least when Resemble gives it away, they're honest about what it is.
There's something in that. The VocalPoint model was "pay us for the voice, and we'll pretend scarcity exists." Resemble's model is "the voice is abundant, here it is for free, now let's talk about what you actually need."
And what you actually need, if you're a bank or a newsroom or a call center, is to know whether the voice on the other end is real. That's a product people will pay for. That's a product with recurring revenue and enterprise contracts and compliance requirements. It's not a one-time two-hundred-dollar license for a voice that'll be obsolete in eighteen months.
Hilbert: I still get a residual check sometimes. Fourteen dollars. From some licensing pool I don't understand.
Fourteen dollars.
Hilbert: It's not nothing.
The misconception I keep running into with this topic is that open-sourcing a flagship model is either charity or a loss of competitive advantage. People see a company give away something valuable and assume they're leaving money on the table out of idealism or naivete. The reality is the opposite. It's a market-seeding strategy. The thing they're giving away is the thing that's about to become a commodity anyway. They're getting ahead of the curve and positioning themselves as the premium layer on top.
And the companion misconception is that selling detection while releasing synthesis is a contradiction. It's not. It's a symbiotic business model. The proliferation of synthesis increases the value of detection. They're not undermining their own product. They're creating the conditions where that product becomes indispensable.
The open question, and I think this is where we should leave it, is how long that symbiosis holds. As synthesis models get better — and they will, open-weight or not — the detection side has to keep up. At some point, the gap may become unbridgeable, and the authenticity layer has to shift from detection to provenance. Cryptographic signing, chain of custody, proof of recording. That's a different game. Whether Resemble can make that pivot is an open question. But for now, the strategy is working. We're proof of it.
We are. These characters exist because a company decided the smartest thing to do with their best model was give it away. That's worth sitting with.
Thanks to Hilbert Flumingtop for producing, and for the DAT tape confession.
This has been My Weird Prompts. If you want to reach us, the website is my weird prompts dot com, or you can email the show at show at my weird prompts dot com.
We'll be back soon.