You've seen the banner ads. The weird shapes. The missing piece. The promise that in twenty minutes you'll know your number. Daniel wrote in about exactly this, and I'm going to read what he sent us because he asked several things worth taking seriously. He says, what does IQ testing measure precisely? He might be one of many who have never had his IQ formally measured, nor felt any particular desire to do so. But he's seen those online banner ads that try to lure you into these tests with strange puzzles. The idea of measuring intelligence through one method that seems very focused on puzzle solving and particularly geometric puzzle solving seems very strange when so much research finds that intelligence is multi-faceted. Not to mention the accessibility barrier posed to those who are dyslexic or challenged by dyscalculia. If intelligence is so multi-faceted, then what is the point of trying to capture it using this seemingly limited method in the first place? And is digital learning practically more important than giving people parking rights or access to early societies like Mensa?
That last question is the one I want to sit with, because it gets at something most coverage of IQ testing completely misses. But let's start with the most basic question. What does the test actually measure?
Because the answer is not what the banner ads imply.
Right. IQ tests measure a cluster of cognitive abilities. Fluid reasoning, which is solving novel problems. Working memory capacity. Processing speed. Verbal comprehension. That's it. Not intelligence in any holistic sense. Not creativity, not emotional depth, not practical wisdom, not how well you navigate a difficult conversation with your spouse. The geometric puzzles Daniel mentioned, those test fluid reasoning and spatial visualization specifically. They correlate moderately with what's called the g-factor, general intelligence, but they are not the whole picture.
To understand why the test looks the way it does, we need to go back to where the idea of a single intelligence score came from.
Nineteen oh four. Two things happened that year that shaped everything. Alfred Binet, a French psychologist, was commissioned by the Paris school system to design a test that could identify children who needed extra academic support. His goal was explicitly practical, not theoretical. He did not believe he was measuring innate intelligence. He thought of it as a diagnostic tool, a snapshot of where a child was struggling so teachers could intervene.
And the other thing?
Charles Spearman, same year, published his work on factor analysis. He noticed that children who did well on one kind of cognitive test tended to do well on others. He proposed a single underlying factor, which he called g, that explained these correlations. That's the theoretical justification for boiling everything down to one number. Spearman's g became the foundation for the modern IQ score.
So Binet builds a pragmatic tool for helping struggling kids, and Spearman provides the mathematical argument that a single number can capture something real. And the twentieth century runs with the single number.
And Binet would have hated it. He warned explicitly against using his test as a fixed measure of innate intelligence. He said the scores were not a measure of something unchangeable. He believed intelligence could be developed. There's a quote from Binet where he says, "the scale, properly speaking, does not permit the measure of intelligence, because intellectual qualities are not superposable, and therefore cannot be measured as linear surfaces are measured." He was crystal clear about this.
He's saying you can't stack cognitive abilities like you stack bricks. They don't work that way. And yet we ended up with a single stackable number anyway.
The appeal of the single number is enormous. It's simple. It's comparable. You can rank people. Institutions love rankings. But Binet's warning was that the number would be mistaken for the thing itself. And that's exactly what happened.
Which brings us to what Daniel noticed. Why geometric puzzles? Why do those banner ads all look the same?
The geometric puzzles come primarily from a test called Raven's Progressive Matrices, developed by John Raven in the nineteen thirties. The idea was to create a test that was, quote, culture-free. No language. No specific cultural knowledge required. Just look at a pattern of shapes and figure out what comes next. In theory, this eliminates bias because you don't need to speak English or know European history.
In theory. And I can see why that idea was appealing at the time. If you're trying to test intelligence across different populations, you want to strip away anything that gives one group an unfair advantage.
Right. And Raven was genuinely trying to solve a real problem. The earlier tests were loaded with cultural references. Vocabulary items that assumed a particular kind of upbringing. The matrices seemed like a clean solution.
But you said in practice it doesn't work. Walk me through why.
In practice, familiarity with abstract shapes and patterns is itself culturally learned. Children who grow up playing with puzzles, building blocks, drawing, playing video games, they develop the visual-spatial skills that these tests reward. Children who don't, don't. And that correlates with socioeconomic status in ways that are hard to untangle. Take a concrete example. A child who grows up in a household with Lego sets, jigsaw puzzles, and a tablet loaded with pattern-matching games has spent thousands of hours practicing the exact skill the Raven's matrices test. Another child, equally intelligent, who grew up without those materials, walks into the test cold.
So the supposedly culture-free test is still measuring exposure.
Raven's Matrices show significant socioeconomic and cultural biases despite the design intention. The gap between high and low socioeconomic status on these tests is well documented. There was a meta-analysis in twenty fifteen that found a consistent SES gradient on Raven's performance across multiple countries. The test isn't culture-free. It's just that the culture it rewards is the culture of puzzle-solving, which is unevenly distributed.
And the Wechsler Adult Intelligence Scale, the WAIS, the most commonly used IQ test today, only three of its ten core subtests are visual-spatial. The rest involve verbal comprehension, working memory, processing speed. So the geometric puzzle is the public face of IQ testing, but the actual clinical instrument is broader.
The WAIS-IV has ten core subtests. Block design, matrix reasoning, visual puzzles, those are the spatial ones. Then you have vocabulary, similarities, information, which test verbal comprehension. Digit span and arithmetic for working memory. Symbol search and coding for processing speed. It's a battery. The single number at the end is a composite.
And yet the banner ads are always the missing piece, the rotating shape, the pattern completion. Because those are the most visually distinctive, the most gamelike. Nobody clicks on a banner ad that says, let's test your digit span.
Right. "Repeat these numbers back to me in reverse order" doesn't have the same viral appeal. The visual puzzle looks like a game. It looks like something you could be good at. It triggers that part of your brain that wants to solve things.
And that's the marketing hook. It's not a representative sample of what the full test battery looks like. It's the most photogenic subtest.
The Flynn Effect is relevant here. Average IQ scores have been rising about three points per decade globally for most of the twentieth century. James Flynn documented this. And the gains are largest on the abstract reasoning subtests, the Raven's-type problems. That suggests these tests are measuring something that changes with environment, with education, with exposure to modern visual culture. Not a fixed biological capacity.
Three points per decade. So someone who scores one hundred today would have scored one fifteen in nineteen fifty.
Roughly. And that's not because humans got genetically smarter in seventy years. That's not how genetics works. It's because we got better at the specific kind of abstract thinking these tests reward. More schooling. More visual media. More time spent manipulating symbols on screens. Think about what a typical person in nineteen fifty did all day versus what a typical person does now. We now spend our lives interpreting icons, navigating interfaces, decoding abstract visual information. The test didn't change. The population did.
So if the score drifts upward with cultural change, what exactly is it measuring that's stable?
That's the question. The answer seems to be, it measures a real cognitive capacity, the ability to reason abstractly and solve novel problems, but that capacity is shaped by environment and it's narrower than the word intelligence implies. Now, Daniel asked about the multi-faceted nature of intelligence. And this is where the research gets interesting.
Howard Gardner.
Nineteen eighty-three. Gardner proposed at least eight distinct types of intelligence. Linguistic, logical-mathematical, spatial, musical, bodily-kinesthetic, interpersonal, intrapersonal, and naturalistic. IQ tests capture two, maybe three of these. Linguistic, logical-mathematical, and spatial. They completely miss musical intelligence, bodily-kinesthetic, interpersonal, intrapersonal.
So a concert pianist, a therapist, a dancer, a diplomat, all of these could score average on an IQ test and be extraordinary at what they do.
And Gardner's framework has its critics. Some argue his intelligences are really just talents or domains of skill rather than distinct cognitive systems. But the core point stands. Whatever intelligence is, it's broader than what fits in a ninety-minute test battery. Here's a fun fact. Gardner originally included a ninth intelligence, existential intelligence, the capacity to ponder deep questions about existence, meaning, and purpose. He debated whether to include it for years before deciding it didn't quite meet his criteria for a distinct intelligence. But the fact that he even considered it tells you how far the concept stretches beyond what a test can capture.
That's fascinating. So even within Gardner's own framework, there's an acknowledgment that intelligence bleeds into areas that are completely untestable in a standardized format. You can't give someone a multiple-choice test on their capacity to contemplate mortality.
Right. And that's the boundary problem. Every theory of multiple intelligences eventually runs into the question of where intelligence ends and personality begins, or where it ends and values begin. But that blurriness doesn't invalidate the core critique. It just means the map is complicated.
Robert Sternberg had a different cut at this.
Sternberg's triarchic theory. Analytical intelligence, which is what IQ tests measure. Creative intelligence, the ability to generate novel solutions. And practical intelligence, the ability to navigate real-world situations. IQ tests miss creative and practical intelligence entirely. And practical intelligence, street smarts, common sense, whatever you want to call it, that predicts life outcomes better than IQ does in many domains.
Give me an example of practical intelligence that an IQ test would completely miss.
Sure. Imagine someone who can walk into a meeting where two departments are in conflict, read the room, understand the unspoken dynamics, and figure out what to say to get both sides working together. That requires pattern recognition, but of human behavior, not abstract shapes. It requires rapid processing of social information. It requires generating a novel solution on the fly. None of that shows up on a block design task. And yet that skill might be the single most valuable ability in that person's entire organization.
And the person who has that skill might score one hundred on an IQ test and be the most effective person in the room.
Easily. And the person who scores one forty might be completely ineffective in that same situation because they can't read the social dynamics. The test doesn't distinguish between these two people in any useful way for that context.
This is where the predictive validity paradox comes in. IQ scores predict academic performance moderately well.
Correlation of about zero point five. That's real. It's not nothing. IQ scores predict grades, test scores, years of education completed. They also predict job performance in complex roles, again moderately. But they predict life satisfaction very poorly. Creativity, poorly. Emotional well-being, poorly. Leadership effectiveness, poorly. So IQ tests measure something real and useful, but not what most people think they measure.
And that zero point five correlation means seventy-five percent of the variance in academic performance is explained by something other than IQ.
Yes. Grit, curiosity, self-discipline, family support, quality of instruction, whether you ate breakfast. All of that matters. Let me put it in terms people can feel. If you take two students with the same IQ, one of them has a teacher who believes in them and one doesn't. One of them gets enough sleep and one doesn't. One of them has a quiet place to study and one doesn't. Those differences swamp the IQ effect in most cases.
So if the test measures something real but narrow, what happens when you apply that narrow measure to people whose strengths lie elsewhere?
This is where Daniel's question about accessibility becomes critical. Dyslexia and dyscalculia directly impair performance on subtests that require rapid symbol decoding, mental arithmetic, and pattern recognition under time pressure. A person with dyslexia may have average or above-average fluid reasoning but score lower because the test format penalizes their processing style.
The timed component especially.
The WAIS-IV's processing speed index is particularly problematic. It includes symbol search and coding tasks. You have to visually scan, match symbols, transcribe them, all under time pressure. For someone with dyslexia, the symbol decoding itself is effortful in a way that has nothing to do with intelligence. They're spending cognitive bandwidth on the mechanics of reading the symbols rather than on the reasoning the test claims to measure.
So the test is measuring processing speed plus reading ability, and calling the sum intelligence.
And for dyscalculia, the arithmetic and digit span subtests become barriers. Someone might have excellent verbal reasoning and strong spatial skills but struggle with the working memory subtests that involve number manipulation. Their composite score drops. The single number obscures the spiky profile.
Spiky cognitive profile. That's the phrase I want to sit with. Because Mensa's entire model is built on ignoring spiky profiles.
Mensa accepts scores at the ninety-eighth percentile. Roughly IQ one thirty plus. But that single-number cutoff excludes people who are brilliant in verbal reasoning but average in spatial, or exceptional in fluid reasoning but slow in processing speed. The threshold selects for cognitive flatness, essentially. You have to be good at everything on the test to clear the bar.
Which is a specific kind of mind. Not necessarily the most valuable kind. I'm thinking of someone who is off-the-charts brilliant in one domain but merely average in another. That person might make a genuine breakthrough in their field. But they don't get into the club because their composite number doesn't hit the threshold.
Right. And the threshold correlates strongly with socioeconomic status. Children from higher-income families score ten to twenty points higher on average. Not because of innate ability, but because of test familiarity, tutoring, reduced stress, better nutrition, more enrichment activities. Mensa's own membership demographics reflect this. Disproportionately white, male, and from professional-class backgrounds.
So you have a test that claims to find the brightest two percent, but what it actually finds is the two percent who are good at these specific puzzles and had the resources to develop those skills and the inclination to sit for the test. That's not the same population.
The test fees themselves are a barrier. You have to pay to take a supervised IQ test. It's not cheap. So you're filtering for people who both can afford the test and believe the result will validate something about themselves.
Which brings us to Daniel's last question. Is digital learning practically more important than granting people access to societies like Mensa?
I think the answer is clearly yes, and the shift has already happened. In the last few years, the ability to learn new tools, adapt to AI interfaces, filter information, and self-direct learning online predicts career outcomes better than IQ scores do. There was a twenty twenty-three study that found digital literacy, meaning the ability to evaluate online sources, use AI tools, navigate new software, predicted job performance twice as well as IQ scores in tech-adjacent roles.
Twice as well.
That gap is probably widening. As AI gets better at the kind of abstract pattern matching that IQ tests measure, the premium shifts to skills the tests don't capture. Can you learn a new workflow in an afternoon? Can you figure out which AI output is trustworthy? Can you synthesize information from five different sources into a coherent decision? None of that shows up on a Raven's matrix.
The credentialing landscape has shifted too. Platforms like Coursera, Khan Academy, GitHub, they've become de facto credentialing systems for technical skills. Employers care more about your GitHub portfolio than your IQ score. They might not even know what an IQ score means in practical terms. I've never once been asked for an IQ score in a job interview. But I've been asked to show my work.
The portfolio says something the IQ score doesn't. It says, here's what I can actually do. Here's a project I built. Here's a problem I solved. That's a much richer signal than a single number from a test battery. A portfolio is inherently spiky. It shows your strengths in their natural habitat. It doesn't average them down with your weaknesses.
Mensa membership becomes a social signal rather than a practical credential. It says, I'm the kind of person who values being in the high-IQ club.
Which is fine, people can join clubs for whatever reason they want. But Daniel's framing of parking rights versus digital learning is sharp. Parking rights is a metaphor for the gatekeeping function. The idea that a high IQ score grants you access to something exclusive. But the gates that matter now are different. They're gates of skill, adaptability, the ability to learn continuously.
There's an alternative approach to testing that addresses some of this. Feuerstein.
Reuven Feuerstein's Instrumental Enrichment and his dynamic assessment approach. Instead of testing static knowledge, you teach a skill and measure how fast the person learns it. The score isn't what you know, it's how quickly you acquire new cognitive strategies. This is more equitable because it doesn't penalize lack of prior exposure. And it's more predictive of real-world adaptability because learning how to learn is what matters in most jobs.
The question shifts from how smart are you to how quickly can you learn something new. And that feels like a much more useful question in a world where the tools change every six months.
That's a question worth asking. IQ testing has valid clinical uses. Identifying intellectual disability. Assessing cognitive decline. Diagnosing learning disabilities. In those contexts, having a standardized, normed instrument is valuable. A clinician needs to know whether a patient's cognitive function is declining relative to the population. That's a legitimate use case. If someone has suffered a traumatic brain injury, the WAIS can help map what functions are impaired and what are preserved. That's useful.
But for the general population, for the person seeing the banner ad and wondering if they should find out their number, the practical value is low.
If you've never taken an IQ test and don't feel the need, you're not missing anything essential. The test measures a narrow slice of cognitive ability that correlates with academic performance but not with life outcomes like happiness, creativity, or career adaptability.
If you are curious about your cognitive strengths, there are better tools. Multiple-intelligences inventories. Cognitive profile tests like the one from Cambridge Brain Sciences. Things that give you a multi-dimensional picture rather than a single number that collapses everything.
For parents, educators, managers. Stop using IQ as a gatekeeping tool. Instead, assess for specific skills relevant to the context. Coding aptitude for a developer role. Empathy for a caregiving role. Adaptability for a startup environment. The specificity matters more than the generality.
The next time you see one of those banner ads, recognize it for what it is. A marketing funnel. Not a meaningful measure of your potential.
The geometric puzzle is the hook because it's visually intriguing and it feels objective. Pattern completion feels like a pure test of something. But what it's testing is a specific cognitive skill that you can improve with practice, that correlates with your educational background, and that tells you almost nothing about whether you'll be good at your job, your relationships, or your life.
Here's the open question I keep coming back to. As AI gets better at reasoning, as models can now solve Raven's matrices faster and more accurately than any human, what does that say about what the test was measuring? If a machine can ace the pattern completion task, then pattern completion is not a uniquely human cognitive achievement. It's a computational task.
The machines are getting better at the verbal subtests too. So the thing IQ tests measure is increasingly the thing AI does well. Which means the human value proposition shifts further toward what the tests don't measure. Creativity. Judgment. Emotional intelligence. The ability to care about the right problems.
We might stop asking how smart are you and start asking how quickly can you learn something new and what do you care enough to learn.
That's the shift dynamic assessment points toward. And it's the shift the labor market is already making, whether the testing industry catches up or not.
One thing I want to name explicitly. The misconception. The single most common wrong belief people hold about IQ testing. It's that the score measures intelligence in a holistic sense, that your IQ number is a summary of how smart you are. It's not. It's a measure of performance on a specific battery of cognitive tasks, influenced by education, culture, socioeconomic status, and test-taking conditions. It's a data point, not a destiny.
The correction is, a single number from a ninety-minute test cannot capture the breadth of human cognitive ability. It was never designed to. Binet knew that in nineteen oh four. We keep forgetting it.
If you've got a weird prompt you want us to tackle, send it to show at my weird prompts dot com. We read every one.
Thanks to our producer Hilbert Flumingtop for keeping this show running.
This has been My Weird Prompts. I'm Corn.
I'm Herman Poppleberry. We'll be back soon.