Four airports in one trip. First one confiscates my water bottle, second one doesn't care about liquids but makes me take off my belt, third one laughs when I ask about the belt and then gets annoyed I left my tablet in the bag, and the fourth one... I don't even know. I just walked through. Nobody asked me anything.
Which is exactly the kind of trip that produces a prompt like this one.
Daniel's been on both sides of the security skepticism conversation with us. He's heard you call it theatre, he's heard the comparisons to how Israel does it. But he wants to push past that debate and get at something more specific. He just did a multi-airport transit, and the inconsistency drove him up the wall. His question, distilled: is this chaos just different security policies, or is it actually a reflection of different scanner technologies at each checkpoint? Where does the millimeter-wave scanner, the one that can supposedly see through clothing, fit into this? And the deeper question. Does cranking up scanner resolution actually make us safer, or does it just multiply the number of weird restrictions and false positives that screeners have to deal with?
That last question is the one that's going to take us somewhere interesting.
So where do we even start? Because the surface answer, the one TSA gives when reporters ask, is that the procedures are the same everywhere but the technology available at each airport may differ. Which sounds reasonable until you actually travel.
Right. And the gap between that statement and the passenger experience is enormous. Same passenger, same day, same airline, different terminals in the same airport. JFK Terminal 4 lets you keep your laptop in the bag and your shoes on. JFK Terminal 1, same airport, requires laptop out and shoes off. That's not a policy difference. That's two different sets of machines at two checkpoints that haven't been upgraded on the same schedule.
So the primary driver here is equipment, not policy.
Almost entirely. And it comes down to three main types of hardware. First, walk-through metal detectors. These are the old arches. They detect metal and nothing else. They can't distinguish a belt buckle from a knife, so belts and shoes come off. Second, millimeter-wave body scanners. These use radio frequency waves to see through clothing, so belts and shoes can stay on, but they have their own problems. Third, CT scanners for carry-on bags. These produce three-dimensional X-ray images that officers can rotate, so laptops and liquids stay in the bag. Older two-dimensional X-ray machines require removal because the image is flat and dense objects obscure what's behind them.
And the rollout of CT scanners is the budget story here. TSA has committed seven hundred and eighty-one million dollars to get these into airports, and as of this year about two hundred and fifty U.S. airports have them. But there are more than four hundred and forty federalized airports. So you've got this patchwork where some checkpoints are upgraded and some aren't, sometimes in the same building.
And that's before you get to the millimeter-wave body scanners versus metal detectors. The machines themselves dictate what you can keep on your body. A metal detector can't tell your belt is harmless. A millimeter-wave scanner can see the shape of the buckle, recognize it's not a threat, and let you through.
So the inconsistency is basically a deployment map. Which checkpoint got the new gear and which didn't.
That's the surface answer. But Daniel's prompt goes deeper, and this is where the physics gets weird. He mentioned that millimeter-wave scanning proves the point that the more data you collect, the more artifacts you collect. He's right, and the reason is water.
Sweat.
Sweat. Millimeter-wave scanners operate in the thirty to three hundred gigahertz range. At those frequencies, the waves penetrate fabric just fine, but they reflect off water. Your skin is wet. You're always perspiring a little. Doug McMakin, the lead researcher who developed this technology at Pacific Northwest National Lab, told ProPublica flat out: the scanner can mistake sweat for a potentially dangerous object.
So the thing the scanner is best at, seeing through clothes to find objects on the body, is also the thing that makes it trip over basic human biology.
And it gets worse when you add the privacy fix. Remember the outcry around twenty ten, twenty eleven? People were horrified that TSA screeners could see detailed images of their bodies. The phrase "virtual strip search" got thrown around a lot.
I remember the images. They were... detailed.
They were. So TSA replaced the detailed body image with a generic unisex outline. A featureless figure, nicknamed the Gumby, that shows only yellow boxes where the machine thinks it's found something. This is called Automated Target Recognition, or ATR software. And it eliminated human discretion.
Wait. The human discretion was the part that was working?
That's the argument Kip Hawley made. He was TSA administrator from two thousand five to two thousand nine. He said trained officers could look at an image and recognize, oh, that's a sweat pattern, that's a breast pocket, that's a clothing fold. The automated system can't. It sees an anomaly and flags it. It doesn't know why.
So the privacy fix made the security worse.
The numbers bear it out. Germany ran a field test at Hamburg Airport, eight hundred and nine thousand passengers. With ATR software, the false positive rate was fifty-four percent. More than half of all alarms were nothing. And thirty-nine percent of all alarms were caused specifically by sweat, buttons, or clothing folds.
Fifty-four percent. That's a coin flip that you're going to get patted down for being damp.
France tested the machines at Charles de Gaulle, eight thousand passengers, and abandoned deployment entirely because of false alarms. Italy's test had a twenty-three percent false alarm rate. The U.S. hasn't published false alarm data since nineteen ninety-six, when it was thirty-one percent at Sea-Tac. And that was with human screeners looking at the images. The early automated software pushed it to thirty-eight and a half percent.
So we don't actually know what the U.S. false alarm rate is right now.
TSA refused to release those numbers to ProPublica in twenty eleven, citing national security. They haven't published updated figures since. We're flying blind on whether the machines are getting better or worse at this.
Which makes Daniel's question about whether higher resolution makes us safer kind of impossible to answer from the data. But the physics points in one direction.
The physics says the tradeoff is inescapable. You pick millimeter waves because they penetrate clothing but reflect off skin. That's the whole point. But water and fabric folds also reflect them. Crank up the resolution, you get more sensitivity to smaller objects, which is what you want for security. You also get more sensitivity to sweat, to seams, to buttons, to the fold of a shirt where it tucks into a waistband. Every one of those looks like something.
And the machine doesn't know the difference between a sheet explosive and a guy who ran to make his connection.
Jan Korte, a German parliament member, called the millimeter-wave scanner a defective product after those test results came out. That's harsh, but you can see where he's coming from. A fifty-four percent false positive rate means the machine is wrong more often than it's right.
There's a comparison that makes this even sharper. The backscatter X-ray scanners, the ones the EU banned over radiation concerns, had a false alarm rate under five percent at Manchester Airport over two and a half million passengers.
And those used ionizing radiation, which is why Europe banned them. So you've got this triangle: you can have low false alarms with radiation exposure, or high false alarms with safe radio waves, or you can skip body scanners entirely and use metal detectors that miss non-metallic threats. Pick your poison.
Daniel also asked about the one that can supposedly see through clothing. Where does that fit?
That is the millimeter-wave scanner. That's its defining capability. The waves go through fabric and bounce off skin and anything on the skin. That's how it works. The "see through clothing" framing was the privacy panic from a decade ago, but it's not wrong. The machine can produce a detailed image of your body. It just doesn't show it to anyone anymore. The Gumby figure is a mask over the raw data.
So the capability is still there. The display is what changed.
And that's where the privacy-security tradeoff really bites. The raw data is detailed enough that a trained human could distinguish a sweat patch from a threat. The automated system can't. So we traded privacy for accuracy, and we may not have gotten much privacy in return, because the machine still collects the data.
I want to sit with that Hamburg number for a second. Thirty-nine percent of alarms were sweat, buttons, or folds. That means four out of ten people who got flagged were flagged because of their clothing or their body temperature. And every one of those flags means a pat-down. A human being puts their hands on another human being because a machine couldn't tell the difference between a zipper and a weapon.
And the pat-down itself is a security event. It takes time, it takes a screener away from watching the line, it creates a confrontation. If half your alarms are false, you're not just wasting time. You're degrading the whole system. The screeners stop trusting the machine.
Which brings us to the second-order question. Does any of this make us safer?
The evidence is... thin. The TSA has a layers-of-security framework, twenty-one layers, from intelligence gathering to hardened cockpit doors. Checkpoint screening is one layer. The argument is that even an imperfect layer contributes to defense in depth.
And Daniel acknowledged that in his prompt. He said screening is an imperfect layer of a stack, and if you believe in defense in depth, all layers play a role. He's not arguing against screening. He's asking why the inconsistency is so extreme.
And whether the inconsistency itself is a problem. There's an argument, and TSA has made versions of this, that unpredictability is a feature. If a would-be attacker can't predict exactly what the screening will look like at a given checkpoint, that's a deterrent.
I've heard that argument. I don't buy it.
Why not?
Because unpredictability as a security strategy only works if the attacker can't adapt. But an attacker can just... watch. They can observe which checkpoints are lax and which are strict. The inconsistency isn't random in a way that confuses adversaries. It's patterned. It's based on which machines are installed where. Anyone doing reconnaissance can map that.
That's fair. And the passenger experience tells you something about how well the layers are working together. If the system were coherent, you'd expect the rules to feel coherent. Instead, you get a guy at one checkpoint confiscating your water and a guy at the next checkpoint waving you through with the same bottle. That erodes trust. And trust matters for security. Passengers who trust the system comply more readily. Passengers who think it's arbitrary start pushing back.
There's also the Israel comparison, which Daniel has raised before and which hangs over this whole discussion.
Ben-Gurion Airport takes a fundamentally different approach. They use profiling and behavioral interviewing as the primary screen. Technology is secondary. Isaac Yeffet, who ran global security for El Al, put it bluntly: technology works well when used to help qualified and well-trained human beings. Technology can never replace the human being. And in the U.S.A., technology is the only security that we have.
Harsh.
Not entirely fair. The U.S. has behavior detection officers, canine teams, air marshals. But the checkpoint is overwhelmingly a technology operation. Israel's checkpoint is overwhelmingly a human operation. They do use body scanners, but after the behavioral assessment, not as the first line.
Ben-Gurion's record is what it is. No successful attack on a plane operating from there since nineteen seventy-two.
But scale matters. Israel screens about fifteen million passengers a year at one primary international airport. The U.S. screened nine hundred and four million passengers last year across more than four hundred and forty airports. Behavioral profiling at that scale would require an enormous workforce. And it raises civil liberties questions the U.S. legal system hasn't resolved.
We're stuck with the machines. Which brings us back to Daniel's core question: are we making the machines better in a way that actually helps, or just in a way that generates more yellow boxes?
There's a Nature Communications paper from twenty twenty-three that points toward one possible path. Researchers demonstrated large-scale single-shot millimeter-wave imaging using sparse antenna arrays and AI reconstruction. They were able to detect concealed centimeter-sized objects with only ten percent of the antenna elements you'd normally need.
AI filling in the gaps.
The idea is that instead of building ever-denser antenna arrays that collect more raw data and generate more artifacts, you collect less data and use AI to reconstruct what's there. The paper showed you can get usable images from a much sparser signal. If that works at scale, it could break the tradeoff. Fewer false alarms because you're not illuminating every sweat droplet, but still enough resolution to find threats.
But AI reconstruction has its own artifact problem. We've talked about this in other contexts. The AI fills in what it expects to see. If it expects to see nothing, it might fill in nothing where there's actually something.
That's the risk. And the paper is a lab demonstration, not a deployed system. Rohde and Schwarz has a newer scanner, the QPS two oh one, that uses AI-powered algorithms and claims improved detection with fewer pat-downs. But we don't have independent false alarm data for it.
The fundamental question, does higher resolution make us safer, is still open. And the inconsistency Daniel experienced is the visible symptom of that unresolved question.
The inconsistency is a deployment map of a technology that hasn't settled. We're in the middle of a transition from metal detectors to millimeter-wave, from two-dimensional X-ray to CT, and different airports are at different points on that curve. The passenger experiences the transition as chaos.
There's one more layer to this that I think matters. The screeners themselves.
The human factor.
We've been talking about machines and policies. But the person in the blue shirt at the checkpoint is making decisions all day. Which alarm to pursue, which to wave through, how thoroughly to pat down. And the machines don't eliminate that discretion. They just change what it's applied to.
The ATR software was supposed to eliminate discretion. The machine flags, the screener responds. But in practice, if the machine is flagging half the passengers and the line is backing up, someone makes a call.
Hilbert: L-three Provision. The first-gen model, before they got the ATR software dialed in. We got three of them in the fall of twenty eleven. I was working the checkpoint at a mid-sized airport, the kind that gets the new equipment in the second wave, after the big hubs.
You were a screener?
Hilbert: Two years. Took the job when the plant closed. The training on the new machines was four hours in a conference room with a PowerPoint. They told us the false alarm rate would be low, the machines would make our jobs easier, and we'd do fewer pat-downs. First week, we did more pat-downs than we'd ever done with the metal detectors. The machines flagged everything. Belt loops, underwire, the seam on the back of a dress shirt. Sweat was the worst. Guy comes in from the parking garage in August in the South, he's glowing like a Christmas tree on the screen.
What did you do?
Hilbert: What they told us not to do. We started ignoring alarms. The supervisor would see the line backing up, thirty people deep, and he'd say just wave them through. Not officially. Never in writing. But if the machine flagged a belt buckle on a businessman for the third time in a row, you stopped patting him down and you let him go.
The inconsistency wasn't just between airports. It was between shifts at the same checkpoint.
Hilbert: Depended who was running the floor. Morning supervisor was strict, afternoon guy just wanted to clear the line. Same machines, same passengers, different experience depending on when you showed up.
That's... that's the whole argument in one anecdote. The technology was supposed to standardize screening. Instead, the false alarm rate was so high that the humans had to override it, and the overrides varied by shift.
Hilbert: The machine gives you a yellow box. It doesn't tell you what to do with it. You still have to decide. And when you've seen the same yellow box on the same belt buckle forty times in a shift, you stop seeing it.
That's a known phenomenon in vigilance tasks. Alert fatigue. The more alarms a system generates, the less attention each alarm receives. A fifty-four percent false positive rate doesn't just waste time. It trains the operator to ignore the system.
Hilbert: I still think about that. How many real threats we might have missed because we were tired of chasing sweat.
The ATR software was supposed to fix that by taking the human out of the loop. But it just moved the human decision to a different point. Instead of deciding what the image shows, you're deciding whether to trust the yellow box.
Hilbert: The machines are better now. I went through a checkpoint last year with one of the newer units and it didn't flag my watch, which the old ones always did. But the core problem hasn't changed. The machine sees something. You have to decide if it's real. And you're making that decision every thirty seconds for eight hours.
The inconsistency Daniel experienced, the different rules at different airports, that's the surface. Underneath it, there's inconsistency between terminals, between shifts, between individual screeners. The whole thing is a stack of human judgments layered on top of machine judgments, and neither layer is perfectly reliable.
Which actually brings us back to Daniel's defense-in-depth point in an unexpected way. He said screening is an imperfect layer in a stack. But what Hilbert's describing is that the screening layer itself is a stack. You've got the machine layer and the human layer, and they don't always agree, and sometimes the human layer overrides the machine layer in ways that aren't documented or consistent.
If the layers within the layer aren't coherent, what does that say about the larger defense-in-depth framework?
It says the framework only works if each layer is doing its job well enough that the next layer can catch what slips through. But if the checkpoint layer is generating noise, false alarms that desensitize the operators, then it's not just failing to catch threats. It's actively making the human layer worse.
The open question Daniel ended on was whether increasing scanner resolution makes us safer or just multiplies the problems. I think the answer, from everything we've walked through, is that it does both. Higher resolution catches smaller threats. It also catches more sweat, more folds, more buttons. And the false alarms aren't just an inconvenience. They degrade the human side of the system.
The Nature Communications paper suggests there might be a way out. Instead of collecting more data and sorting through more artifacts, collect less data and reconstruct intelligently. But that's a lab result, not a product. And even if it works, you still have the human at the end of the chain making judgment calls.
Daniel's prompt started with a travel story, four airports, four sets of rules. The answer is that the rules aren't rules. They're the visible output of a machine deployment schedule, a physics tradeoff that nobody has solved, and a human system that's constantly negotiating with its own technology. The inconsistency isn't a bug in an otherwise coherent system. It is the system.
The question of whether we're safer... we don't have the data to answer it. TSA hasn't published false alarm rates in three decades. We're running a massive security operation and we can't say whether the machines at the center of it are getting better or worse.
That's probably where we should leave it. Not with an answer, but with the shape of the problem. Thanks to Daniel for the prompt, and thanks to Hilbert Flumingtop for producing and for... well, for reminding us that the PowerPoint training was four hours.
The show was produced by Hilbert Flumingtop. This has been My Weird Prompts. If you've got a weird prompt of your own, email us at show at my weird prompts dot com.
We'll be back soon.