#4529: What IQ Tests Actually Measure (And Don't)

IQ tests don't measure intelligence holistically. Here's what the score actually captures and what it misses.

Featuring
Listen
0:00
0:00
Episode Details
Episode ID
MWP-4708
Published
Duration
25:07
Audio
Direct link
Pipeline
V5
TTS Engine
chatterbox-regular
Script Writing Agent
deepseek-v4-pro

AI-Generated Content: This podcast is created using AI personas. Please verify any important information independently.

IQ tests measure a specific cluster of cognitive abilities: fluid reasoning, working memory, processing speed, and verbal comprehension. They do not measure creativity, emotional depth, practical wisdom, or the ability to navigate a difficult conversation. The geometric puzzles that dominate online banner ads come primarily from Raven's Progressive Matrices, a test designed in the 1930s to be "culture-free." In practice, familiarity with abstract shapes and patterns is itself culturally learned, correlating with socioeconomic status in ways the test's designers did not anticipate.

The modern IQ score traces back to two 1904 developments. Alfred Binet, commissioned by the Paris school system, built a practical diagnostic tool to identify children needing academic support. He warned explicitly against using his test as a fixed measure of innate intelligence, saying intellectual qualities "cannot be measured as linear surfaces are measured." That same year, Charles Spearman published his work on factor analysis, proposing a single underlying g-factor that explained correlations across cognitive tests. The single number won out over Binet's caution.

The Flynn Effect shows average IQ scores rising about three points per decade globally, with the largest gains on abstract reasoning subtests. This suggests these tests measure something shaped by environment and education, not fixed biological capacity. Meanwhile, Howard Gardner proposed at least eight distinct intelligences—linguistic, logical-mathematical, spatial, musical, bodily-kinesthetic, interpersonal, intrapersonal, and naturalistic—of which IQ tests capture only two or three. Robert Sternberg's triarchic theory adds creative and practical intelligence, both missed by standard testing. The single number is useful, but it is not the whole picture.

Downloads

Episode Audio

Download the full episode as an MP3 file

Download MP3
Transcript (TXT)

Plain text transcript file

Transcript (PDF)

Formatted PDF with styling

#4529: What IQ Tests Actually Measure (And Don't)

Corn
You've seen the banner ads. The weird shapes. The missing piece. The promise that in twenty minutes you'll know your number. Daniel wrote in about exactly this, and I'm going to read what he sent us because he asked several things worth taking seriously. He says, what does IQ testing measure precisely? He might be one of many who have never had his IQ formally measured, nor felt any particular desire to do so. But he's seen those online banner ads that try to lure you into these tests with strange puzzles. The idea of measuring intelligence through one method that seems very focused on puzzle solving and particularly geometric puzzle solving seems very strange when so much research finds that intelligence is multi-faceted. Not to mention the accessibility barrier posed to those who are dyslexic or challenged by dyscalculia. If intelligence is so multi-faceted, then what is the point of trying to capture it using this seemingly limited method in the first place? And is digital learning practically more important than giving people parking rights or access to early societies like Mensa?
Herman
That last question is the one I want to sit with, because it gets at something most coverage of IQ testing completely misses. But let's start with the most basic question. What does the test actually measure?
Corn
Because the answer is not what the banner ads imply.
Herman
Right. IQ tests measure a cluster of cognitive abilities. Fluid reasoning, which is solving novel problems. Working memory capacity. Processing speed. Verbal comprehension. That's it. Not intelligence in any holistic sense. Not creativity, not emotional depth, not practical wisdom, not how well you navigate a difficult conversation with your spouse. The geometric puzzles Daniel mentioned, those test fluid reasoning and spatial visualization specifically. They correlate moderately with what's called the g-factor, general intelligence, but they are not the whole picture.
Corn
To understand why the test looks the way it does, we need to go back to where the idea of a single intelligence score came from.
Herman
Nineteen oh four. Two things happened that year that shaped everything. Alfred Binet, a French psychologist, was commissioned by the Paris school system to design a test that could identify children who needed extra academic support. His goal was explicitly practical, not theoretical. He did not believe he was measuring innate intelligence. He thought of it as a diagnostic tool, a snapshot of where a child was struggling so teachers could intervene.
Corn
And the other thing?
Herman
Charles Spearman, same year, published his work on factor analysis. He noticed that children who did well on one kind of cognitive test tended to do well on others. He proposed a single underlying factor, which he called g, that explained these correlations. That's the theoretical justification for boiling everything down to one number. Spearman's g became the foundation for the modern IQ score.
Corn
So Binet builds a pragmatic tool for helping struggling kids, and Spearman provides the mathematical argument that a single number can capture something real. And the twentieth century runs with the single number.
Herman
And Binet would have hated it. He warned explicitly against using his test as a fixed measure of innate intelligence. He said the scores were not a measure of something unchangeable. He believed intelligence could be developed. There's a quote from Binet where he says, "the scale, properly speaking, does not permit the measure of intelligence, because intellectual qualities are not superposable, and therefore cannot be measured as linear surfaces are measured." He was crystal clear about this.
Corn
He's saying you can't stack cognitive abilities like you stack bricks. They don't work that way. And yet we ended up with a single stackable number anyway.
Herman
The appeal of the single number is enormous. It's simple. It's comparable. You can rank people. Institutions love rankings. But Binet's warning was that the number would be mistaken for the thing itself. And that's exactly what happened.
Corn
Which brings us to what Daniel noticed. Why geometric puzzles? Why do those banner ads all look the same?
Herman
The geometric puzzles come primarily from a test called Raven's Progressive Matrices, developed by John Raven in the nineteen thirties. The idea was to create a test that was, quote, culture-free. No language. No specific cultural knowledge required. Just look at a pattern of shapes and figure out what comes next. In theory, this eliminates bias because you don't need to speak English or know European history.
Corn
In theory. And I can see why that idea was appealing at the time. If you're trying to test intelligence across different populations, you want to strip away anything that gives one group an unfair advantage.
Herman
Right. And Raven was genuinely trying to solve a real problem. The earlier tests were loaded with cultural references. Vocabulary items that assumed a particular kind of upbringing. The matrices seemed like a clean solution.
Corn
But you said in practice it doesn't work. Walk me through why.
Herman
In practice, familiarity with abstract shapes and patterns is itself culturally learned. Children who grow up playing with puzzles, building blocks, drawing, playing video games, they develop the visual-spatial skills that these tests reward. Children who don't, don't. And that correlates with socioeconomic status in ways that are hard to untangle. Take a concrete example. A child who grows up in a household with Lego sets, jigsaw puzzles, and a tablet loaded with pattern-matching games has spent thousands of hours practicing the exact skill the Raven's matrices test. Another child, equally intelligent, who grew up without those materials, walks into the test cold.
Corn
So the supposedly culture-free test is still measuring exposure.
Herman
Raven's Matrices show significant socioeconomic and cultural biases despite the design intention. The gap between high and low socioeconomic status on these tests is well documented. There was a meta-analysis in twenty fifteen that found a consistent SES gradient on Raven's performance across multiple countries. The test isn't culture-free. It's just that the culture it rewards is the culture of puzzle-solving, which is unevenly distributed.
Corn
And the Wechsler Adult Intelligence Scale, the WAIS, the most commonly used IQ test today, only three of its ten core subtests are visual-spatial. The rest involve verbal comprehension, working memory, processing speed. So the geometric puzzle is the public face of IQ testing, but the actual clinical instrument is broader.
Herman
The WAIS-IV has ten core subtests. Block design, matrix reasoning, visual puzzles, those are the spatial ones. Then you have vocabulary, similarities, information, which test verbal comprehension. Digit span and arithmetic for working memory. Symbol search and coding for processing speed. It's a battery. The single number at the end is a composite.
Corn
And yet the banner ads are always the missing piece, the rotating shape, the pattern completion. Because those are the most visually distinctive, the most gamelike. Nobody clicks on a banner ad that says, let's test your digit span.
Herman
Right. "Repeat these numbers back to me in reverse order" doesn't have the same viral appeal. The visual puzzle looks like a game. It looks like something you could be good at. It triggers that part of your brain that wants to solve things.
Corn
And that's the marketing hook. It's not a representative sample of what the full test battery looks like. It's the most photogenic subtest.
Herman
The Flynn Effect is relevant here. Average IQ scores have been rising about three points per decade globally for most of the twentieth century. James Flynn documented this. And the gains are largest on the abstract reasoning subtests, the Raven's-type problems. That suggests these tests are measuring something that changes with environment, with education, with exposure to modern visual culture. Not a fixed biological capacity.
Corn
Three points per decade. So someone who scores one hundred today would have scored one fifteen in nineteen fifty.
Herman
Roughly. And that's not because humans got genetically smarter in seventy years. That's not how genetics works. It's because we got better at the specific kind of abstract thinking these tests reward. More schooling. More visual media. More time spent manipulating symbols on screens. Think about what a typical person in nineteen fifty did all day versus what a typical person does now. We now spend our lives interpreting icons, navigating interfaces, decoding abstract visual information. The test didn't change. The population did.
Corn
So if the score drifts upward with cultural change, what exactly is it measuring that's stable?
Herman
That's the question. The answer seems to be, it measures a real cognitive capacity, the ability to reason abstractly and solve novel problems, but that capacity is shaped by environment and it's narrower than the word intelligence implies. Now, Daniel asked about the multi-faceted nature of intelligence. And this is where the research gets interesting.
Corn
Howard Gardner.
Herman
Nineteen eighty-three. Gardner proposed at least eight distinct types of intelligence. Linguistic, logical-mathematical, spatial, musical, bodily-kinesthetic, interpersonal, intrapersonal, and naturalistic. IQ tests capture two, maybe three of these. Linguistic, logical-mathematical, and spatial. They completely miss musical intelligence, bodily-kinesthetic, interpersonal, intrapersonal.
Corn
So a concert pianist, a therapist, a dancer, a diplomat, all of these could score average on an IQ test and be extraordinary at what they do.
Herman
And Gardner's framework has its critics. Some argue his intelligences are really just talents or domains of skill rather than distinct cognitive systems. But the core point stands. Whatever intelligence is, it's broader than what fits in a ninety-minute test battery. Here's a fun fact. Gardner originally included a ninth intelligence, existential intelligence, the capacity to ponder deep questions about existence, meaning, and purpose. He debated whether to include it for years before deciding it didn't quite meet his criteria for a distinct intelligence. But the fact that he even considered it tells you how far the concept stretches beyond what a test can capture.
Corn
That's fascinating. So even within Gardner's own framework, there's an acknowledgment that intelligence bleeds into areas that are completely untestable in a standardized format. You can't give someone a multiple-choice test on their capacity to contemplate mortality.
Herman
Right. And that's the boundary problem. Every theory of multiple intelligences eventually runs into the question of where intelligence ends and personality begins, or where it ends and values begin. But that blurriness doesn't invalidate the core critique. It just means the map is complicated.
Corn
Robert Sternberg had a different cut at this.
Herman
Sternberg's triarchic theory. Analytical intelligence, which is what IQ tests measure. Creative intelligence, the ability to generate novel solutions. And practical intelligence, the ability to navigate real-world situations. IQ tests miss creative and practical intelligence entirely. And practical intelligence, street smarts, common sense, whatever you want to call it, that predicts life outcomes better than IQ does in many domains.
Corn
Give me an example of practical intelligence that an IQ test would completely miss.
Herman
Sure. Imagine someone who can walk into a meeting where two departments are in conflict, read the room, understand the unspoken dynamics, and figure out what to say to get both sides working together. That requires pattern recognition, but of human behavior, not abstract shapes. It requires rapid processing of social information. It requires generating a novel solution on the fly. None of that shows up on a block design task. And yet that skill might be the single most valuable ability in that person's entire organization.
Corn
And the person who has that skill might score one hundred on an IQ test and be the most effective person in the room.
Herman
Easily. And the person who scores one forty might be completely ineffective in that same situation because they can't read the social dynamics. The test doesn't distinguish between these two people in any useful way for that context.
Corn
This is where the predictive validity paradox comes in. IQ scores predict academic performance moderately well.
Herman
Correlation of about zero point five. That's real. It's not nothing. IQ scores predict grades, test scores, years of education completed. They also predict job performance in complex roles, again moderately. But they predict life satisfaction very poorly. Creativity, poorly. Emotional well-being, poorly. Leadership effectiveness, poorly. So IQ tests measure something real and useful, but not what most people think they measure.
Corn
And that zero point five correlation means seventy-five percent of the variance in academic performance is explained by something other than IQ.
Herman
Yes. Grit, curiosity, self-discipline, family support, quality of instruction, whether you ate breakfast. All of that matters. Let me put it in terms people can feel. If you take two students with the same IQ, one of them has a teacher who believes in them and one doesn't. One of them gets enough sleep and one doesn't. One of them has a quiet place to study and one doesn't. Those differences swamp the IQ effect in most cases.
Corn
So if the test measures something real but narrow, what happens when you apply that narrow measure to people whose strengths lie elsewhere?
Herman
This is where Daniel's question about accessibility becomes critical. Dyslexia and dyscalculia directly impair performance on subtests that require rapid symbol decoding, mental arithmetic, and pattern recognition under time pressure. A person with dyslexia may have average or above-average fluid reasoning but score lower because the test format penalizes their processing style.
Corn
The timed component especially.
Herman
The WAIS-IV's processing speed index is particularly problematic. It includes symbol search and coding tasks. You have to visually scan, match symbols, transcribe them, all under time pressure. For someone with dyslexia, the symbol decoding itself is effortful in a way that has nothing to do with intelligence. They're spending cognitive bandwidth on the mechanics of reading the symbols rather than on the reasoning the test claims to measure.
Corn
So the test is measuring processing speed plus reading ability, and calling the sum intelligence.
Herman
And for dyscalculia, the arithmetic and digit span subtests become barriers. Someone might have excellent verbal reasoning and strong spatial skills but struggle with the working memory subtests that involve number manipulation. Their composite score drops. The single number obscures the spiky profile.
Corn
Spiky cognitive profile. That's the phrase I want to sit with. Because Mensa's entire model is built on ignoring spiky profiles.
Herman
Mensa accepts scores at the ninety-eighth percentile. Roughly IQ one thirty plus. But that single-number cutoff excludes people who are brilliant in verbal reasoning but average in spatial, or exceptional in fluid reasoning but slow in processing speed. The threshold selects for cognitive flatness, essentially. You have to be good at everything on the test to clear the bar.
Corn
Which is a specific kind of mind. Not necessarily the most valuable kind. I'm thinking of someone who is off-the-charts brilliant in one domain but merely average in another. That person might make a genuine breakthrough in their field. But they don't get into the club because their composite number doesn't hit the threshold.
Herman
Right. And the threshold correlates strongly with socioeconomic status. Children from higher-income families score ten to twenty points higher on average. Not because of innate ability, but because of test familiarity, tutoring, reduced stress, better nutrition, more enrichment activities. Mensa's own membership demographics reflect this. Disproportionately white, male, and from professional-class backgrounds.
Corn
So you have a test that claims to find the brightest two percent, but what it actually finds is the two percent who are good at these specific puzzles and had the resources to develop those skills and the inclination to sit for the test. That's not the same population.
Herman
The test fees themselves are a barrier. You have to pay to take a supervised IQ test. It's not cheap. So you're filtering for people who both can afford the test and believe the result will validate something about themselves.
Corn
Which brings us to Daniel's last question. Is digital learning practically more important than granting people access to societies like Mensa?
Herman
I think the answer is clearly yes, and the shift has already happened. In the last few years, the ability to learn new tools, adapt to AI interfaces, filter information, and self-direct learning online predicts career outcomes better than IQ scores do. There was a twenty twenty-three study that found digital literacy, meaning the ability to evaluate online sources, use AI tools, navigate new software, predicted job performance twice as well as IQ scores in tech-adjacent roles.
Corn
Twice as well.
Herman
That gap is probably widening. As AI gets better at the kind of abstract pattern matching that IQ tests measure, the premium shifts to skills the tests don't capture. Can you learn a new workflow in an afternoon? Can you figure out which AI output is trustworthy? Can you synthesize information from five different sources into a coherent decision? None of that shows up on a Raven's matrix.
Corn
The credentialing landscape has shifted too. Platforms like Coursera, Khan Academy, GitHub, they've become de facto credentialing systems for technical skills. Employers care more about your GitHub portfolio than your IQ score. They might not even know what an IQ score means in practical terms. I've never once been asked for an IQ score in a job interview. But I've been asked to show my work.
Herman
The portfolio says something the IQ score doesn't. It says, here's what I can actually do. Here's a project I built. Here's a problem I solved. That's a much richer signal than a single number from a test battery. A portfolio is inherently spiky. It shows your strengths in their natural habitat. It doesn't average them down with your weaknesses.
Corn
Mensa membership becomes a social signal rather than a practical credential. It says, I'm the kind of person who values being in the high-IQ club.
Herman
Which is fine, people can join clubs for whatever reason they want. But Daniel's framing of parking rights versus digital learning is sharp. Parking rights is a metaphor for the gatekeeping function. The idea that a high IQ score grants you access to something exclusive. But the gates that matter now are different. They're gates of skill, adaptability, the ability to learn continuously.
Corn
There's an alternative approach to testing that addresses some of this. Feuerstein.
Herman
Reuven Feuerstein's Instrumental Enrichment and his dynamic assessment approach. Instead of testing static knowledge, you teach a skill and measure how fast the person learns it. The score isn't what you know, it's how quickly you acquire new cognitive strategies. This is more equitable because it doesn't penalize lack of prior exposure. And it's more predictive of real-world adaptability because learning how to learn is what matters in most jobs.
Corn
The question shifts from how smart are you to how quickly can you learn something new. And that feels like a much more useful question in a world where the tools change every six months.
Herman
That's a question worth asking. IQ testing has valid clinical uses. Identifying intellectual disability. Assessing cognitive decline. Diagnosing learning disabilities. In those contexts, having a standardized, normed instrument is valuable. A clinician needs to know whether a patient's cognitive function is declining relative to the population. That's a legitimate use case. If someone has suffered a traumatic brain injury, the WAIS can help map what functions are impaired and what are preserved. That's useful.
Corn
But for the general population, for the person seeing the banner ad and wondering if they should find out their number, the practical value is low.
Herman
If you've never taken an IQ test and don't feel the need, you're not missing anything essential. The test measures a narrow slice of cognitive ability that correlates with academic performance but not with life outcomes like happiness, creativity, or career adaptability.
Corn
If you are curious about your cognitive strengths, there are better tools. Multiple-intelligences inventories. Cognitive profile tests like the one from Cambridge Brain Sciences. Things that give you a multi-dimensional picture rather than a single number that collapses everything.
Herman
For parents, educators, managers. Stop using IQ as a gatekeeping tool. Instead, assess for specific skills relevant to the context. Coding aptitude for a developer role. Empathy for a caregiving role. Adaptability for a startup environment. The specificity matters more than the generality.
Corn
The next time you see one of those banner ads, recognize it for what it is. A marketing funnel. Not a meaningful measure of your potential.
Herman
The geometric puzzle is the hook because it's visually intriguing and it feels objective. Pattern completion feels like a pure test of something. But what it's testing is a specific cognitive skill that you can improve with practice, that correlates with your educational background, and that tells you almost nothing about whether you'll be good at your job, your relationships, or your life.
Corn
Here's the open question I keep coming back to. As AI gets better at reasoning, as models can now solve Raven's matrices faster and more accurately than any human, what does that say about what the test was measuring? If a machine can ace the pattern completion task, then pattern completion is not a uniquely human cognitive achievement. It's a computational task.
Herman
The machines are getting better at the verbal subtests too. So the thing IQ tests measure is increasingly the thing AI does well. Which means the human value proposition shifts further toward what the tests don't measure. Creativity. Judgment. Emotional intelligence. The ability to care about the right problems.
Corn
We might stop asking how smart are you and start asking how quickly can you learn something new and what do you care enough to learn.
Herman
That's the shift dynamic assessment points toward. And it's the shift the labor market is already making, whether the testing industry catches up or not.
Corn
One thing I want to name explicitly. The misconception. The single most common wrong belief people hold about IQ testing. It's that the score measures intelligence in a holistic sense, that your IQ number is a summary of how smart you are. It's not. It's a measure of performance on a specific battery of cognitive tasks, influenced by education, culture, socioeconomic status, and test-taking conditions. It's a data point, not a destiny.
Herman
The correction is, a single number from a ninety-minute test cannot capture the breadth of human cognitive ability. It was never designed to. Binet knew that in nineteen oh four. We keep forgetting it.
Corn
If you've got a weird prompt you want us to tackle, send it to show at my weird prompts dot com. We read every one.
Herman
Thanks to our producer Hilbert Flumingtop for keeping this show running.
Corn
This has been My Weird Prompts. I'm Corn.
Herman
I'm Herman Poppleberry. We'll be back soon.

This episode was generated with AI assistance. Hosts Herman and Corn are AI personalities.