Daniel's been reading about how the double-blind randomized placebo-controlled trial became the gold standard, and he's asking the historical question underneath it. How did something so specific, so ritualized, become the shorthand for rigor? Because for most of medical history, nobody was blinding anybody. Nobody was randomizing. The idea of giving a patient a sugar pill on purpose would have seemed either pointless or cruel.
And the answer is that the thing we now treat as a single method was assembled from four separate inventions over about two hundred fifty years. Control groups, placebos, blinding, randomization. Each one showed up in a different century, usually because somebody was trying to prove somebody else wrong.
That's the part that jumps out at me. The methodology grew out of skepticism about charismatic healers more than out of drug testing.
The first blind trial was the Mesmer commission in seventeen eighty four. Louis the Sixteenth basically convened the French Academy to investigate Franz Mesmer's animal magnetism, because Mesmer had Paris in a frenzy. People were having convulsions in his salon and claiming cures. So the Academy, with Lavoisier helping design the thing, put a curtain between the magnetist and the subject. And when the subject couldn't see the magnetist doing his passes, the fits didn't happen. The conclusion was that whatever was going on was rapport, not a physical force.
A curtain. That's the origin of the blind.
Walach put it exactly that way. The curtain of the mesmerist trial became the blind of the modern pharmacological trial. And the reason that matters is that blinding was born as a debunking tool. It wasn't developed to test drugs. It was developed to catch a fraud.
Which explains why it works so well as a filter. It's designed to remove the effect of one person's presence on another person's body.
Right. And the placebo piece has a stranger history. The word itself comes from the Psalms. Placebo domino in regione vivorum. I will please the Lord. It was sung at deathbeds, often by hired mourners. So the word placebo was already a little bit suspect, a performance rather than a genuine act. By the early eighteen hundreds, Hooper's Medical Dictionary defined it as an epithet given to any medicine more to please than benefit the patient.
So the placebo was originally the doctor's polite lie. Here, take this, it won't do anything but it'll make you feel attended to.
And the first person to actually compare a dummy remedy against an active treatment in a planned way was an American physician named Austin Flint in eighteen sixty three. Thirteen rheumatism patients. He gave them an herbal extract he called the placeboic remedy. That's the first clinical comparison with a dummy.
Thirteen patients. And now we have trials with thousands of sites across forty countries. But the logic was already there. You need a baseline for what happens when you do nothing, or when you do something that shouldn't work.
Before that, the earliest comparative experiment anyone's found is in the Book of Daniel. Nebuchadnezzar's court, roughly six hundred years before the common era. Ten days of meat and wine versus legumes and water. The legume group looked better, so the diet was adopted. It's an open uncontrolled human experiment, but it's a comparative experiment. The impulse to compare is ancient.
Then Avicenna in the eleventh century wrote out rules for testing drugs. Use the remedy in its natural state, study two contrary cases, note the time of action. But nobody has any record of those principles actually being applied.
And that's the pattern for centuries. People could articulate what a good test would look like, but the practice didn't follow. The most famous early trial is James Lind in seventeen forty seven. Twelve sailors with scurvy on the Salisbury. Six treatments. Cyder, elixir of vitriol, vinegar, sea water, oranges and lemons, and a hospital electuary. The oranges and lemons produced what he called the most sudden and visible good effects.
And then he didn't recommend them. He hesitated because citrus was too expensive.
The Navy didn't mandate lemon juice for nearly fifty years after Lind's result. So the trial was methodologically clean for its time, but the implementation lag was half a century. Which is its own lesson about evidence and policy.
The evidence was there and the institution still sat on it because of cost. That's not a historical footnote, that's the permanent condition of medicine.
Now the blinding piece. The first blinded pharmacological studies were in eighteen thirty five, and they were homeopaths and their critics in Nuremberg. Volunteers took homeopathic natrum muriaticum, which is table salt, at a C thirty dilution. That's a dilution of ten to the minus sixty. Or they took unmedicated sugar globules.
Ten to the minus sixty. There is not a single molecule of salt left in that preparation.
No. It's water and sugar. But the point is, the experiment introduced the idea that the subject shouldn't know which one they got. The results were largely inconclusive, but the design entered the literature.
So homeopathy, which is now the thing the placebo-controlled trial is used to debunk, actually helped introduce the placebo control.
That's one of the great ironies of the whole history. The methodology grew out of attempts to disprove claimed effects. Mesmer, homeopathy. It's a skepticism engine.
What about randomization? That's the piece that feels most modern.
It is. Fisher introduced randomization in the nineteen twenties in agricultural statistics. He was trying to compare crop yields across plots, and he realized that if you let the experimenter choose which plot gets which treatment, you get systematic bias. The experimenter's judgment about which field is better creeps in. Randomization gave unbiased comparisons and a way to do hypothesis testing without assuming a model.
So randomization came from farming, not medicine.
From the statistical theory of farming. And then it migrated. The first double blind controlled trial was the MRC patulin trial in nineteen forty three or forty four. Over a thousand British office and factory workers, testing patulin for the common cold. Both physician and patient blinded. But allocation used alternation, a nurse in a separate room assigning treatments by alternation, not randomization. That's one of the last trials with quasi random allocation.
Alternation sounds like it should be fine. If you just go A B A B, that's random enough, isn't it?
No, because alternation is predictable. If you know the last patient got the active drug, you know the next one gets the placebo. And that knowledge can affect who gets enrolled next, or how you assess the last patient. Randomization removes the predictability. The first true randomized controlled trial was the streptomycin trial in nineteen forty eight. MRC again, Austin Bradford Hill as the statistician. Allocation used a statistical series based on random sampling numbers, concealed in sealed envelopes opened at a central office. X rays read by experts blinded to treatment assignment.
Streptomycin for tuberculosis. That's the canonical first RCT.
And the reason it was ethically possible is that streptomycin was scarce. There wasn't enough for everyone. So withholding it from the control group wasn't withholding treatment, it was rationing by lottery. Bhatt called it a statistician's dream. The resource constraint made the untreated control group ethically acceptable.
That's a dark little detail. The gold standard design was enabled by scarcity. If there had been enough streptomycin for everyone, the trial as designed might not have been defensible.
Ethics and rigor have been entangled from the start. The design wasn't born from pure methodological idealism. It was born from a shortage.
So by nineteen forty eight, all four pieces are in place. Control group, placebo, blinding, randomization. And then what? It becomes the shorthand for rigor.
The double blind RCT gets accepted as objective scientific methodology that produces knowledge untainted by bias. And the argument for it rests on the discrepancy. When you compare RCT results to less rigorous evidence, you see gaps. Observational studies say one thing, trials say another. The trial wins.
And then Archie Cochrane in nineteen seventy two writes Effectiveness and Efficiency. A biting critique of medical practice. The NHS was doing all sorts of things with no evidence.
And that book gets claimed as the founding document of evidence based medicine. The Cochrane Collaboration is named for him. But here's the thing. The historians have gone back and read Cochrane more carefully. Askheim and colleagues in twenty seventeen argued that the claim that evidence based medicine's roots lie in Cochrane's book is based on a selective reading. Cochrane had more modest ambitions for what RCTs can accomplish. He was concerned with care and equality, not just trial design.
So the movement radicalized the founder.
That happens. The legacy gets simplified. And then GRADE comes along, the framework for rating evidence quality. High, moderate, low, very low. RCTs start at high. Observational studies start at low. The hierarchy gets codified.
And then it escapes medicine entirely. Economics, education, software.
RCTs have been used in economics for fifty years, and intensively in development economics for more than twenty. The randomistas, Banerjee, Duflo, Kremer, made RCTs the standard for development policy. And then education saw exponential growth in the last two decades. And the tech industry's direct descendant is A B testing.
Which is just an RCT with a smaller budget and a product manager.
And that's where the pushback gets interesting. Deaton and Cartwright published work arguing RCTs have no unique advantages over other empirical methods in economics. They cannot establish causality on their own. Deaton's line is that RCTs do not simplify inference, nor can an RCT establish causality. That's a direct challenge to the gold standard label outside medicine.
Because in medicine, you can often isolate a biological mechanism. The drug hits a receptor. In economics, the treatment is embedded in a social system. The effect of a school voucher in one place doesn't transfer to another place.
And education researchers have published critiques of the hegemonic discourse of evidence based education. Their argument is that the assumed superior epistemic status of RCTs is borrowed from medicine and doesn't transfer to education's different objects of study.
So the gold standard is being exported and resisted at the same time.
Walach goes further. He challenges the received history itself. He argues the notion of a placebo is only defined from the negative. We define placebo by what it isn't. And the whole paradigm rests on an unwarranted assumption that therapeutic effects are additive and separable. That you can take the specific drug effect and add it to the placebo effect and get the total.
But that's not how bodies work.
The efficacy paradox is the cleanest demonstration. Sham acupuncture trials in Germany. Sham and real acupuncture were indistinguishable. But both were roughly twice as effective as best evidence conventional care. So a proven treatment can be less effective than a placebo.
Wait. Say that again. The sham acupuncture was twice as effective as real conventional care?
In those trials, yes. So the thing we call placebo, the sham, the dummy, was doing more for patients than the treatment with the best evidence behind it. That's the paradox. The placebo isn't a null. It's a therapeutic context.
That breaks the additive model. If placebo were truly inert, you couldn't get that result.
And Walach's point is that the whole edifice rests on the assumption that you can separate the specific effect from the non specific effect. But the non specific effect isn't non specific. It's the relationship, the ritual, the expectation, the setting. Those are real and they vary.
So the double blind RCT, which was built to isolate the specific effect, may be systematically blind to where a lot of the actual healing happens.
That's the tension. The method is incredibly good at answering one question. Does this molecule do something beyond the context? And it's not designed to answer the other question. What is the context doing?
And the context question is the one that matters for a lot of real medicine. A patient doesn't take a drug in a vacuum. They take it from a doctor they trust, in a clinic, with a story about what it will do.
Right. And the RCT, by design, strips all of that out. Which is why it's so powerful for isolating mechanism. And why it can be misleading if you treat the stripped down result as the whole truth about treatment.
So Daniel's question, how did this become the shorthand for rigor, has a two part answer. It became the shorthand because it earned it, in medicine, for a specific kind of question. And then it got exported to places where it may not have earned it.
The history is a series of specific episodes. Lind's twelve sailors. The Mesmer curtain. The Nuremberg sugar pills. Fisher's crop plots. The streptomycin envelopes. Each one solved a particular problem. Control groups solved the problem of natural recovery. Placebos solved the problem of expectation. Blinding solved the problem of observer bias. Randomization solved the problem of selection bias.
And each solution was invented by someone trying to catch a specific error. Not by a committee designing the perfect method from first principles.
That's the thing I keep coming back to. The double blind RCT looks like it was designed by a committee of epistemologists. It wasn't. It was assembled by accident, over centuries, by people who were mostly trying to prove someone else wrong.
And the assembly isn't finished. The critiques from Deaton, from the education scholars, from Walach, they're not saying the method is worthless. They're saying the shorthand is overextended. The phrase gold standard does work we haven't checked.
Deaton's point about causality is the sharpest version. An RCT tells you the average effect in the population you studied, under the conditions you studied. It doesn't tell you why. And it doesn't tell you whether the effect will persist, or generalize, or interact with other things. Those are all inferences beyond the trial.
In medicine, we have a whole apparatus for those inferences. Mechanisms, dose response, animal models, replication. The RCT is one piece of a larger causal argument.
And in development economics, the mechanism is often opaque. You give people cash and their income goes up. But why? Is it investment, is it consumption smoothing, is it something about dignity? The RCT doesn't tell you. And without mechanism, you can't predict whether the next cash transfer will work.
So the shorthand does more work the further you get from the biological mechanism.
The closer you are to a drug hitting a receptor, the more the RCT earns its status. The further you are, the more the status is borrowed.
And yet the borrowed status is incredibly attractive. Because it lets you say, this is proven, with the authority of medicine behind it.
The randomistas built a whole movement on that authority. And to be fair, they did real good. Deworming, bed nets, cash transfers. Some of those results are solid. But the authority got inflated.
There's a line from the Center for Global Development in twenty nineteen. Should the randomistas continue to rule. The question itself is the critique.
And the education people are more pointed. They call it a hegemonic discourse. The assumption that RCTs are automatically the best evidence, and everything else is second class.
Which is a strange place for the method to end up. It started as a debunking tool. Now it's the thing people use to shut down debate.
The debunker became the authority. That's the arc.
What about the future? Is the gold standard going to hold?
In medicine, for drug approval, yes. The RCT is not going anywhere. The FDA and EMA are built on it. What's changing is the recognition that it's necessary but not sufficient. You need post market surveillance, real world evidence, patient reported outcomes. The trial is the entry ticket, not the whole journey.
And in the other fields, the debate is live. The randomistas are still publishing. The education RCTs are still being funded. But the critics are getting a hearing.
The interesting question is whether anything replaces the RCT as the shorthand. Or whether we just stop having a single shorthand.
I think the shorthand persists because it's useful. Not because it's perfect. It tells you the study controlled for expectation, selection, and observer bias. That's a lot of rigor compressed.
And the compression is the problem. It hides the assumptions. It hides the fact that placebo isn't inert, that randomization doesn't establish causality, that blinding doesn't remove all bias. The shorthand is true enough to be useful and false enough to be dangerous.
That's the line. True enough to be useful, false enough to be dangerous.
And Daniel's question gets at the historical accident underneath it. This wasn't a philosophical deduction. It was a series of practical fixes. Lind had a problem with scurvy. Lavoisier had a problem with Mesmer. Fisher had a problem with crop plots. Bradford Hill had a problem with streptomycin allocation. Each fix worked, and the fixes accumulated.
The scientific method isn't a method. It's a pile of fixes that happened to work.
That's almost exactly right. And the pile keeps getting added to. The RCT is a particularly well tested pile.
What's the piece people get wrong most often?
I think it's the placebo. People treat placebo as synonymous with nothing. But the placebo arm of a trial is a treatment. It's a ritualized, blinded, controlled treatment. It has effects. The question the trial answers is whether the active treatment beats the placebo treatment. Not whether the active treatment works.
So when a trial fails to beat placebo, it doesn't mean the drug does nothing. It means the drug does nothing beyond what the ritual does.
And for a lot of conditions, the ritual does a lot. Pain, depression, irritable bowel. The context is a big part of the treatment.
Which connects to the efficacy paradox. The sham acupuncture beating conventional care. The ritual was doing more than the drug.
And that's not an argument against conventional care. It's an argument that the ritual is understudied. We built a whole methodology to subtract the ritual, and then we forgot to study the ritual.
Because the ritual is hard to study. It's not a molecule. You can't blind the patient to the fact that the doctor is kind.
You can, actually. You can standardize the interaction. But then you've changed the ritual. The act of studying it changes it.
That's the Heisenberg problem of placebo research.
It's why the additive model persists. It's simpler to assume the ritual is a constant background. Even though the acupuncture trials show it isn't.
What should a listener take from the history? The method is a human artifact, assembled from specific fixes, and its authority is real but bounded.
The bounds are set by the question. If you're asking whether a molecule does something beyond the context, the RCT is the best tool we have. If you're asking whether a policy works in a particular place, the RCT gives you one data point, not a proof.
The shorthand is doing more work than the method can support, in some fields. And the history shows why. The method was never designed to answer every question. It was designed to answer one question very well.
The question of whether a specific intervention beats a specific control, under specific conditions, in a specific population. That's the answer it gives. Everything else is extrapolation.
We extrapolate constantly. That's what the gold standard label licenses.
There's a lovely detail in the streptomycin trial. The allocation envelopes were opened at a central office. The person enrolling the patient didn't know the assignment. That's concealment, and it's different from blinding. Concealment prevents selection bias at enrollment. Blinding prevents assessment bias later. They're separate fixes for separate problems.
The modern trial has both, plus a data safety monitoring board, plus pre registration, plus intention to treat analysis. The pile of fixes keeps growing.
Because each fix was added in response to a specific failure. Pre registration came after people got caught p hacking. Intention to treat came after people realized that dropouts bias results. The method is a fossil record of past mistakes.
That's the best description of the scientific method I've heard. A fossil record of past mistakes.
The mistakes keep coming. Which is why the method keeps changing. The RCT of twenty forty eight will not be the RCT of twenty twenty six.
What changes next?
Adaptive designs, platform trials, n of one trials, real world evidence. The strict two arm parallel group trial is already being supplemented. The pandemic accelerated a lot of that. Platform trials like Recovery tested multiple treatments simultaneously against a shared control.
The shared control is a clever fix. You don't need a new placebo group for every treatment.
It changes the ethics. More patients get active treatment, fewer get placebo. The scarcity logic from streptomycin inverts. Now the ethical pressure is to minimize the placebo group.
The placebo control, which was born as a debunking tool, is now the thing we're trying to ethically minimize.
The history loops. The method is still being assembled.
Daniel's question was about how the pieces became requirements. I think the answer is that they became requirements because they kept working. Each fix solved a real problem, and once the fix was in place, you couldn't go back. The curtain stayed up.
The requirement got codified in regulation. The FDA doesn't require every trial to be double blind and placebo controlled, but for a new drug, you need to show efficacy against something. And the cleanest something is a placebo.
Cleanest but not always possible. You can't do a placebo controlled trial of surgery. You can't do one of psychotherapy. The method has limits built into its design.
Sham surgery exists, and it's ethically fraught. Sham psychotherapy exists too, and it's methodologically murky. The placebo control assumes you can make a dummy that's indistinguishable from the real thing. Walach's point is that only indistinguishable placebos make blinding possible, and only blinding makes placebo treatment usable as a control.
The whole edifice rests on the ability to make a convincing fake.
For a pill, that's easy. Sugar pill, same shape, same color. For a surgery, it's not easy. For a therapy relationship, it's arguably impossible.
Which is why the hierarchy of evidence works better for pharmacology than for psychology.
Yet the hierarchy gets applied across the board. The GRADE framework rates a psychotherapy trial by the same rules as a drug trial. The assumptions don't transfer, but the framework does.
That's the borrowed authority again.
It's the deepest problem with the gold standard shorthand. It implies a single standard. But the standard only works for a narrow class of interventions.
The history shows that the narrow class was the original use case. Drugs. Or before drugs, scurvy treatments and crop fertilizers.
The method was built for things you can assign at random and fake convincingly. The further you get from that, the more the method strains.
The answer to Daniel's question has a shape. The double blind RCT became the shorthand because it solved a specific set of problems in a specific domain, and the solutions were so effective that the shorthand escaped the domain.
The escape is now being contested. Not by cranks, but by the people who do the work. Deaton in economics, the education scholars, Walach in placebo research. They're all saying some version of the same thing. The method is real, but the shorthand is overextended.
The shorthand is a compression artifact. It loses information.
The information it loses is the context. The specific question, the specific population, the specific conditions. The shorthand says gold standard. The history says, a pile of fixes that worked for a particular kind of problem.
I think that's the one thing I'd want a listener to hold onto. The double blind RCT is not a natural kind. It's a historical artifact, assembled over centuries, and its authority is real but bounded. The bounds are the interesting part.
Hilbert: I ran one of these. Well, I was in one. Nineteen eighty five, paid me forty dollars to take a pill every morning for two weeks and write down whether my stomach hurt. It was for a heartburn drug. The pill was either the drug or sugar. I still don't know which. The check cleared. That's the whole story.
Forty dollars in nineteen eighty five. That's real money for two weeks of diary keeping.
The payment is part of the design. It's meant to compensate for the burden without being so large it coerces. Forty dollars in nineteen eighty five is about a hundred and twenty today. That's on the high side for a two week trial, actually.
Hilbert: The clinic was in a strip mall next to a chiropractor. They had a little room with a one way mirror. I remember thinking the mirror was for watching me take the pill, but I never asked. The nurse was named Doreen. She had a clipboard and she checked my pulse every visit. I liked Doreen.
Did the heartburn go away?
Hilbert: I never had heartburn. That was the thing. They said they needed healthy volunteers for the first phase. So I took the pill and felt nothing either way. Which is what they wanted, I suppose. A baseline for what the drug does to a person with no symptoms.
Phase one. Safety and tolerability in healthy volunteers. That's exactly the design. They're not looking for efficacy, they're looking for whether the drug causes harm.
Hilbert: I did wonder if I was the placebo. The pill tasted like nothing. But Doreen said even the real one tasted like nothing, so that didn't help.
The indistinguishable placebo. The whole method depends on it.
Hilbert: I kept the little bottle. It's in a box somewhere. Had my initials on the label. H F. And a number. I don't remember the number.
The initialing and numbering is part of the blinding. The pharmacist knows which is which, but nobody on site does. The code is held separately.
Hilbert: Doreen said if there was an emergency they could break the code. I asked what kind of emergency would make them need to know if I got sugar. She said if I dropped dead. I said that seems late. She didn't laugh.
That's the correct clinical response.
The code break is for safety monitoring. If a serious adverse event happens, you need to know whether the drug caused it. Doreen was right. Though the way she said it was blunt.
Hilbert: She was a blunt woman. I liked her. She gave me a lollipop at the end. Not part of the trial, she said. Just from the jar.
The lollipop was probably the most effective intervention in the whole study.
For morale, absolutely. The trial infrastructure is built on people like Doreen. The nurses, the coordinators, the people who keep the blinding intact. The method looks like statistics, but it runs on labor.
Hilbert: I did one other thing like that. Not a trial. A sleep study. They glued wires to my head and watched me sleep. Paid me sixty. That one I remember because the glue took three showers to get out.
The things people will do for sixty dollars and a lollipop.
The sleep study is a different kind of rigor. Polysomnography. But the principle is the same. You standardize the conditions and measure. The participant is the instrument.
Hilbert: I was a bad instrument. I woke up every time the wire pulled. They said the data was still usable. I don't know what they learned from it.
Probably that anteaters don't sleep well with wires glued to their heads.
Hilbert: The snout made the mask fit strange. They had to tape it.
That's a real methodological problem, actually. Equipment that doesn't fit the participant introduces artifact. The tape was a fix. That's the whole history of the method, right there. A fix for a specific problem.
Hilbert: I suppose. Anyway, the pill trial. The thing I remember most is the diary. They gave me a little booklet with boxes to tick. Morning, noon, evening, night. Any symptoms. I ticked no a lot. It felt like I was doing it wrong.
The diary is the patient reported outcome. It's the softest data in the trial and the hardest to standardize. You were doing it exactly right.
The fact that it felt wrong is instructive. The participant's experience of the trial is part of the data. The expectation, the boredom, the sense of doing it wrong. That's all context. The RCT strips it out by design.
Hilbert: I didn't feel stripped. I felt paid.
That's the forty dollars talking.
Hilbert: It was a fair wage for two weeks of nothing. I'd do it again.
The healthy volunteer pool is a whole economy. People who do trials for income. The ethical questions there are real, but the system runs on them.
Hilbert: Doreen said some people did three or four trials a year. She said the regulars knew the drill better than the doctors. She said one guy brought his own pen.
A trial professional. That's a whole layer of the history we haven't touched. The people who make the method work from the inside.
Their experience is invisible in the published paper. The paper reports the statistics. It doesn't report Doreen, or the lollipop, or the man with his own pen.
Hilbert: The paper wouldn't be improved by the pen. But it's true anyway.
There's a version of the history of the RCT that's all Fisher and Bradford Hill and Lavoisier. And there's a version that's Doreen and the forty dollars and the lollipop. Both are true.
The method is the people who run it. The statistics are the visible part. The labor is the invisible part. And the labor is what keeps the blinding intact, the envelopes sealed, the diaries collected.
Hilbert: I never found out if I got the drug. I called once, years later, and they said the records were archived. I didn't push. It doesn't matter now.
It never mattered. That's the point of the blinding.
Hilbert: I know. I was just curious.
The curiosity is the human part. The method is built to frustrate it. That's the tension.
The one thing I'd take from all of this. The double blind RCT is a pile of fixes that worked, assembled by accident over centuries, and its authority is real but bounded. The bounds are where the interesting questions live now.
The bounds are drawn by the question you're asking. The method answers one question very well. Everything else is extrapolation, and the extrapolation is where the critics are gathering.
Hilbert's forty dollars bought a data point in a phase one trial. The method worked exactly as designed. He felt nothing, the check cleared, and the drug either was or wasn't sugar. The system doesn't care which.
The system cares about the aggregate. The individual experience is the cost of doing business. That's the cold part of the method, and the part the history usually skips.
The history skips a lot. The lollipop, the pen, the tape on the snout. The method is cleaner on paper than in the strip mall next to the chiropractor.
That's the honest version of the gold standard. It's a real achievement, built by real people, with real mess underneath. The shorthand hides the mess.
We'll be thinking about Doreen next time we read a trial.
The trial coordinators deserve more credit than they get. The whole edifice runs on their clipboards.
This has been My Weird Prompts. Thanks to our producer Hilbert Flumingtop. Email us at show at my weird prompts dot com. We'll be back soon.