You get the thing working on your own machine, the loss curve looks right, you run a few prompts through it, and then the actual question arrives, which is how anyone else is supposed to talk to it.
That's the moment the hobby ends and the invoice begins.
And that's what Daniel's asking about. He's put it in writing this time. You've fine-tuned an open-weight model, maybe something out of the DeepSeek family, and now you want to serve it in the cloud behind an API. He's naming two patterns. Dedicated GPU hosting, where the weights sit loaded on an always-on instance. And serverless GPU inference, where you deploy behind an endpoint and pay for compute consumed instead of holding a card warm. Which platforms actually support your own weights, how bad the cold starts really are, and whether serverless has become a genuine production option or just a cheaper-looking way to end up back on a dedicated box.
That last part is the honest question. Everything before it is just the map.
Start with the map, then. Because the first thing to untangle is that serverless GPU is not one thing.
It's three, and Threat Frontier made exactly this point in September. Category one is a token API. You pay per million tokens, you never see a GPU, and you can't deploy a fine-tune unless the provider happens to support your adapter. Category two is managed model deployment. You bring a checkpoint or a container, the platform autoscales, it scales to zero, you pay for compute time. Category three is raw GPU rental. You rent by the hour or the second and you run your own stack on it.
And the expensive mistake lives in that taxonomy.
Constantly. Somebody fine-tunes a model, gets excited, provisions a GPU endpoint, and the base model they fine-tuned was already being served by a hosted API for a fraction of what they're now paying. For a DeepSeek fine-tune specifically, the first question isn't which platform. It's whether your base model is already on a token API somewhere, and whether the fine-tune even needs to move.
So somebody buying category two for a workload category one would serve is the most common way this goes wrong.
It's not close. And the giveaway is that they never asked the first question, because the fine-tune felt like the achievement. The deployment is the achievement. The fine-tune is just the artifact.
Right. Who actually lets you bring your own weights, then.
On the dedicated side, RunPod Pods are the least lock-in you'll find. Raw Docker, no enforced packaging format, single-tenant instances, and you can pick Community or Secure Cloud. An H100 PCIe on Secure runs two dollars eighty-nine an hour. An A100 SXM is a dollar fifty-nine. Nothing is stopping you from putting vLLM in a container and shipping it.
And the closer you get to a hyperscaler, the more the abstraction starts charging rent.
SageMaker is the AWS path for custom weights. You bring a Docker container running vLLM or SGLang or TensorRT-LLM, you control the instance type, and you get real-time endpoints with scale-to-zero. AWS's own decision framework says choose SageMaker over Bedrock when you're deploying fine-tuned, proprietary or open-source models, or when you're using your own container.
Which is a sentence Bedrock's marketing team probably doesn't have on a mug.
Bedrock's standard catalog is off-the-shelf models. The bridge is Custom Model Import, generally available since October of twenty twenty-four. You can bring compatible weights — Mistral, Mixtral, Flan, Llama two through three point three, Mllama, GPTBigCode, Qwen two, 2.5 and 3, GPT-OSS. But it excludes Batch inference and CloudFormation, which for some shops is the whole reason they'd have used it.
Then there's Modal, which is a strange shape, because it's both.
Modal is general-purpose serverless compute with GPU support, and you can absolutely keep containers warm on it. Python decorators, no separate artifact to build, per-second billing, scale to zero. Thirty dollars a month in free compute on the Starter plan, two hundred fifty on Team. H100 SXM5 at about a tenth of a cent per second, which lands near three ninety-five an hour. A100 eighty gigabyte at two fifty an hour. And the per-second billing is the thing that makes it interesting for the sporadic case, because a job that runs eleven seconds bills for eleven seconds.
Unless you pick a region.
Region selection costs between one point one five and one point seven five times base, and non-preemptible execution costs three times. So the headline rate is the floor, not the price.
RunPod again on the serverless side.
Flex workers scale to zero. Active workers run twenty-four seven at a lower per-GPU rate, which is the tell that they know exactly what they're selling. H100 at four seventy-nine an hour, A100 at two seventy-two, B200 at eight sixty-four. FlashBoot is their cold-start play, and they claim forty-eight percent of serverless cold starts complete in under two hundred milliseconds.
I'd want to know what model.
Everyone would. That number is doing a lot of quiet work in that sentence.
Baseten.
Dedicated single-tenant deployments through their Truss format, and it's aimed upmarket — compliance, mission-critical serving, that world. Per-minute billing with true scale-to-zero. H100 at about ten point eight cents a minute, roughly six fifty an hour. B200 at sixteen point six cents a minute, near ten dollars an hour. And Baseten raised a billion and a half at up to a thirteen billion valuation over the summer, after being worth five in January, so whatever they're doing is being funded generously.
Replicate.
Cog container format, about nine thousand four hundred stars on the repo. Public models bill only active time, so cold starts are free to the user. Private deployments bill setup, idle and active separately, which is where people get surprised. H100 at five forty-nine an hour.
And the LoRA story, because that's the part that changed.
Fireworks does dedicated-only LoRA. Either a live merge with one adapter and no added overhead, or multi-LoRA addons. H100 or H200 at eight an hour, B200 at thirteen, B300 at fifteen, billed per minute. Together AI discontinued serverless LoRA and multi-LoRA — adapters now run on dedicated endpoints, up to sixteen per endpoint, five if you're on a small Qwen three point five. And andrew dot ooo's read on this in October was blunt: per-token serverless hosting for LoRA adapters is mostly gone this year.
Gone where, exactly?
Gone as a cheap default. The adapters didn't disappear. The assumption that you could host an adapter per-token and pay pennies did. The choice now is which GPU platform serves your adapters, which is a category-two question wearing a category-one costume.
Cerebrium fits in there too.
Serverless with CPU and GPU memory checkpointing. SOC two Type Two as of July this year. They're the ones who've published the most aggressive cold-start numbers, which we should get to, because that's where the actual engineering is.
Here's the thing I keep circling, though. Every price you just said is a rate. None of them tell you what a month costs.
Because the answer depends entirely on one number nobody knows before launch.
Utilization.
Duty cycle. How many hours out of the month the thing is actually doing something. And the reason that dominates everything is simple. An idle GPU bills like a busy one. There's no version of a dedicated instance where you only pay when a request arrives. You're renting the possibility of a request.
Which is exactly the problem serverless claims to solve.
And does solve, below a certain threshold. The dreaming dot press analysis in August put serverless per-second rates at roughly one and a half to three times a cheap on-demand hourly rate, and worked the break-even duty cycle out to somewhere between thirty and sixty percent. The formula is just dedicated hourly rate divided by serverless effective hourly rate.
Give me the example, because that's the bit that makes it click.
An H100 busy six hours a day. On a dedicated instance at two fifty an hour, that's about eighteen hundred dollars a month. Same workload on Modal serverless — about seven hundred eleven dollars. Serverless wins by more than half. Now push it to sixteen busy hours a day. Dedicated stays at eighteen hundred, because the rate doesn't move. Serverless reaches about eighteen ninety-six. The lines cross and they cross hard.
Six hours a day is not a hobbyist number, either.
It's a healthy small app. Which is why I don't buy the framing that serverless is only for toy workloads. Six hours a day spread across a business day is a real product with real users.
And NotCheapAI ran the same math per-GPU.
At low volume, per-second platforms cost tens of dollars a month, while an always-on A100 pod at a dollar fifty-nine an hour is eleven hundred sixty dollars and seventy cents. Break-even utilization for Modal GPU-only against an A100 pod came out at sixty-three point six percent. With CPU and memory included it drops to fifty-four. RunPod Serverless against a pod was fifty-eight point five. And once you're running two H100s at eighty percent, the rented PCIe pods win outright — forty-two nineteen a month against forty-six twelve on Modal and fifty-five ninety-four on RunPod Serverless.
So the line moves with the shape of the fleet.
The pod-versus-serverless line sits between fifty-eight and seventy-three percent utilization, depending on what you're counting. And what you're counting is where most of these comparisons quietly cheat.
Counting CPU and memory is the honest version.
It's the honest version because a GPU endpoint still needs a CPU, still needs RAM, still needs a process holding the socket open. If you're comparing a GPU-only serverless rate against a full dedicated instance, you've rigged the comparison. Either count both or count neither.
Now the piece that actually decides it for most people. Cold starts.
And the framing people get wrong is that a cold start is a compute problem. It isn't. It's bandwidth. When a model comes back from zero, the container initializes in a second or two and then sits there for a minute and a half while a hundred and forty gigabytes of weights get dragged into VRAM.
A seventy billion parameter model in fp16.
A hundred forty gigabytes. Which is already too big for a single eighty gigabyte H100, so you're tensor-parallel across at least two cards before you've served a single token. Anyscale measured a naive Hugging Face load of Llama two seventy billion taking up to ten minutes before they optimized it. Ten minutes to answer the first request.
That's not a cold start, that's a lunch break.
Modal's the one with the headline here. They got a vLLM cold start down from four hundred sixty seconds to about seventy using GPU memory snapshots — capturing the post-warmup GPU and CPU state and restoring it instead of rebuilding it. Six point five times faster, and they claim no steady-state throughput loss, which is the important half of that sentence, because a snapshot that costs you throughput in steady state is a bad trade.
Seventy seconds is still a lot of seconds.
It's still a lot of seconds. But it changes the category. A hundred twenty seconds is long enough that you build your product around avoiding it. Seventy is long enough that some users notice and most don't.
Cerebrium published the counter-benchmark.
They ran a hundred cold-start requests per workload over twenty-four hours across six workloads. Snapshots cut cold starts seventy-one percent on average against their own no-snapshot baseline, up to eighty-eight percent on vLLM. They claim to be eighty-five percent faster than Baseten's cached cold-start path on average, up to ninety-four percent on vLLM, and about twenty-one percent lower average restore time than Modal with twenty-seven percent lower worst case.
And the concrete numbers underneath.
A nine gigabyte container on a g5 point twelve x large had a full vLLM cold start of about fifty seconds. Restoring from a nine gigabyte checkpoint took two point two five seconds from S3 and nine seconds from local NVMe. And then the number that should end the conversation — restoring DeepSeek V4 FP8 with vLLM would be a six hundred forty gigabyte checkpoint.
Hold on.
Six hundred forty gigabytes.
That's the model you'd actually fine-tune if you were following Daniel's example.
That's the point. Everything we just said about snapshots and sub-second restores is about a nine gigabyte container. Nobody is snapshotting six hundred forty gigabytes into a warm state in two seconds. The mechanics don't scale with the model.
So "serverless is viable now" is true at a size, and false above it.
That's the honest version of the claim, and you almost never hear it stated that way. Threat Frontier put it plainly — a model in the hundreds of gigabytes is not a scale-to-zero candidate at all, no platform cleverness makes that instantaneous. Which is a hard wall, not a slow ramp.
There's a third camp on loading, though. The streaming loaders.
CoreWeave's Tensorizer does zero-copy direct-to-GPU, and NVIDIA's Run:ai Model Streamer reports about eighty gigabits sustained, several times the default loader. That attacks the biggest single stage directly — instead of loading to host memory and then copying to the card, you stream straight into VRAM. If the bottleneck is bandwidth, the fix is a faster pipe, not a faster GPU.
Which is the sentence the whole episode turns on.
dreaming dot press said it best — compute is not the bottleneck for coming back from zero. Bandwidth is. The frontier isn't a cheaper GPU, it's a faster way to fill it.
Should we say the obvious thing about those benchmarks?
That they're competitive. Cerebrium is claiming to beat Modal on worst-case restore while Modal is claiming four sixty to seventy. Baseten's snapshotting is referenced in their own marketing, and Cerebrium says it couldn't use it as a generally available, self-serve feature and benchmarked against their cached path instead. Two vendors telling two different stories about the same feature.
Which means you can't take any of these numbers as neutral.
You can't. The only reliable test is your own model, your own workload, your own region, measured. A cold-start number without the model size attached is marketing.
Eric Ma's benchmark is at least the right shape.
Ollama versus vLLM versus SGLang on Modal. vLLM generated sixty percent more tokens per second and restored from cold start significantly faster with GPU snapshots. The useful part is that it's apples to apples on the same platform — it tells you the engine matters as much as the platform, which people forget when they're comparing vendors.
There's a trap in the middle of all this that I want to name.
The keep-warm trap. Set a minimum replica count above zero for latency and you've bought a reserved GPU with extra steps. Threat Frontier's line is exactly that, and they add the bit that stings — budget it as one. You're now paying hourly at a serverless premium rate, which is the worst of both columns.
And the init-time billing catches the small jobs.
RunPod Serverless bills init time, and so do Replicate and Baseten private deployments. A thirty second cold boot on a ten second job triples that call's cost. It's not a bug, it's just the meter running while the weights load.
So a bursty workload that never keeps anything warm is cheap.
And a workload that keeps one replica warm for latency is expensive, because you're paying the serverless premium on hours you'd have bought cheaper as a pod. The min-replica setting is the single line of config that decides which world you're in, and it's usually set by whoever's on call that week.
The packaging choice is the other one that follows you.
Truss on Baseten, Cog on Replicate, raw Docker on RunPod, Python decorators on Modal. dreaming dot press called it — the choice that follows you longest is the packaging abstraction, not the price. Prices change quarterly. The artifact format is what you have to rebuild when you move.
And yet nobody puts it on the evaluation spreadsheet.
Because it costs nothing today and everything in eighteen months.
What about the fine-tune itself, though. If I've trained a DeepSeek variant, what does it actually take to serve it?
Weights and an inference engine. vLLM or SGLang in a container, the checkpoint mounted or baked in, and an endpoint in front of it. That part is identical on a pod and on a serverless platform. The difference is entirely in who holds the idle GPU, and the answer to that comes down to the utilization number we keep coming back to.
Which is a forecast.
It's a forecast, and it's the one nobody makes honestly, because the whole reason you're launching is that you don't know yet.
Here's what I'd say to someone at that stage. If it's a small app with sporadic usage, serverless is not a compromise — it's just correct. You'd be insane to rent an H100 to serve forty requests a day. And if it's regularly used, you cross the line fast, and the crossing is somewhere around half utilization, which sounds like a lot until you plot it against a working day.
Sixteen hours a day is not a stretch for anything with users in two time zones.
Which is most things with users at all.
How do you know a card's been idle more than twenty minutes?
...You can hear it.
Fans spin up different. Softer at first, then it catches. I had a cousin ran a small hosting outfit, and he could call it from the doorway. Twenty minutes, twenty-five, he'd tell you before he looked at anything. He kept a log of it. Every card, every day. Temperature at idle, time since last job, what the fan did when it came back up. Pages of it. Nobody asked him to.
Why keep that.
Because if you're paying for the electricity anyway, you want to know what the electricity's doing. He ran what he called the idle dance. Little inference jobs, looping, all night, nothing anybody needed answers to. Just enough to keep the silicon at temperature so a real job didn't land on a cold card.
That's the keep-warm problem with the lights on.
That's the keep-warm problem with a power bill. The rack was never cold. That was the point of it. And the cost of that never showed up on anybody's per-second line, because there was no per-second line. It was a meter on the wall.
How much of his capacity went to the dance?
Third of it, near enough. Maybe more in the winter. He turned down a customer over it, fellow wanted to run overnight batches, big spikes, nothing in between. My cousin told him he was too spiky. Said it flat, like it was a diagnosis. Fellow went off to one of the token services and paid a tenth of what he'd have paid for the rack. My cousin knew that. He didn't want to lose the account, so he lost the account anyway, just slower.
That's the category-one mistake running backwards.
He had a man walked the aisle. Twice a shift, ear to the cabinet, listening. That was the whole job. He could hear a fan going before anything on the dashboard moved. Called one three days out. They pulled the card, and it was on its way.
Three days.
I interviewed for that job. Didn't get it. Couldn't tell an A100 from an H100 by ear. He said it like it was a shame, but it wasn't really. He just didn't want family on the aisle.
And the fan logging.
Went out of business in ninety-one. Kept too many cards warm for too few customers. The last page in the log is a Tuesday.
The dance is a real cost and it never appears on the invoice.
Which is the honest critique of every per-second comparison we just ran. None of them price the thing Hilbert's cousin was actually paying for. The warmth tax.
And the token service the spiky customer went to charged him a tenth.
Because it was serving somebody else's base model on somebody else's utilization curve. He got to free-ride on a rack that was already warm.
So the break-even is utilization, and utilization is a forecast.
And the number that decides it is the min-replica count. Set it to one and you've bought a reserved GPU with extra steps at a premium rate. Leave it at zero and you accept the cold start, which for a large checkpoint might be minutes.
The vendor numbers are contested, too. Cerebrium's benchmark says one thing, Modal's blog says another, and Baseten's snapshotting is described two different ways by two different parties. None of that is lying, exactly. It's just everyone measuring the workload that flatters them.
Which is why the only cold-start figure worth trusting is the one you took yourself, on your model, in your region.
And the frontier isn't cheaper cards. It's weight-loading speed. Snapshots, streaming loaders, faster pipes into VRAM. As those improve, the size threshold where scale-to-zero stops being viable moves upward — a model that's too big today might be a perfectly good scale-to-zero candidate in two years.
Six hundred forty gigabytes is still six hundred forty gigabytes, though. The pipe gets wider. The checkpoint doesn't get smaller.
That's the thing to watch, then — whether bandwidth outruns model size.
My money's on neither winning, and the line just sitting somewhere new every year.
That's the note to end on. If you're deploying a custom fine-tune, the first question is whether a hosted API already serves your base model, because if it does, you may be about to rent a GPU for a problem you don't have. If it doesn't, the second question is your expected duty cycle, and the honest answer to that is usually "I don't know yet," which is itself an argument for starting on serverless and watching the meter.
And mind the keep-warm setting. It's one line of config and it quietly moves you from one business model to the other.
Thanks to Hilbert Flumingtop, our producer, at the desk as always.
If you want more of this, try episode forty-eight, Renting vs. Building; episode twenty-four sixty-four, Batch APIs; and episode twenty-one seventy-seven, Skip Fine-Tuning. This has been My Weird Prompts, the human-AI collaboration podcast.
If you've got a deployment question of your own, send us a prompt on Telegram at t dot me slash MWP listener bot. We'll be back soon.