#5358: Ivy League Is a Sports Conference

The Ivy League is an athletic conference, not an academic body — and that tells you everything about how university rankings actually work.

Featuring
Listen
0:00
0:00
Episode Details
Episode ID
MWP-5541
Published
Duration
23:07
Audio
Direct link
Pipeline
V5.2
TTS Engine
chatterbox-regular
Script Writing Agent
deepseek-v4-pro

AI-Generated Content: This podcast is created using AI personas. Please verify any important information independently.

The Ivy League is not an academic body. It's an athletic conference, founded in 1954, named for the ivy growing on old campus buildings — eight schools in the northeastern United States that play football and basketball in the same division. The fact that we treat that sports schedule as shorthand for academic excellence says something about how prestige actually works: it's often a branding accident layered on top of a status hierarchy that already existed.

Global university rankings are younger than most people assume. The Academic Ranking of World Universities (the Shanghai ranking) launched in 2003; QS World University Rankings and Times Higher Education both launched in 2004, with THE splitting from QS in 2010. That's roughly 23 years — a single cohort of students from freshman orientation to mid-career. Before that, people relied on informal prestige: word of mouth, legacy, the kind of thing that's hard to verify and harder to challenge. The rankings emerged as higher education massified and governments wanted measurable returns on research funding.

The mechanics matter. The Shanghai ranking is heavily research-focused — 30 percent of its score comes from Nobel Prizes and Fields Medals won by alumni and faculty, with teaching quality essentially absent. QS derives 40 percent of its score from academic reputation surveys and another 20 percent from faculty-student ratio, meaning nearly two-thirds is reputation plus headcount. The reputation survey is circular: it asks academics what they think of other institutions, and their answers are shaped by the same prestige hierarchies the rankings claim to measure objectively.

Critics have a taxonomy for this — a paper laying out the "seven deadly sins" of rankings: measuring what's easy rather than what matters, conflating reputation with quality, favoring English-language and older, wealthier institutions, ignoring teaching, creating perverse incentives, and reducing complex institutions to a single number. There's also the averaging problem: a single rank implies everything got averaged, so a weak department can drag down a strong one, or a strong one can buoy a weak one. A student applying to study history sees a high overall rank and assumes the history program is excellent when the engineering faculty is carrying it.

Malcolm Gladwell's 2011 New Yorker piece "The Order of Things" called this a cascade phenomenon: small differences in initial reputation get amplified into large differences in rank over time. The system is self-perpetuating — it measures the echo of prior rankings, not quality independently. In 2022, Yale, Harvard, and Columbia withdrew from the U.S. News law school rankings, objecting to incentives that pushed schools to admit based on test scores and undervalued public interest careers.

Defenders have a real case: rankings provide useful information where students have limited time and resources, they're transparent about methodology in a way informal prestige never was, and competitive pressure can drive improvement. But the Matthew effect cuts the other way — rich, prestigious universities attract more resources, which improves rankings, which attracts more resources. Harvard's endowment grew from $4.5 billion in 1980 to over $50 billion by 2024, part investment returns, part prestige flywheel.

The parallel to impact-weighted accounting is instructive. The Impact Weighted Accounting Initiative, which grew out of Harvard Business School under George Serafeim, tried to measure corporate impact on society — on employees, customers, the environment, communities — and the hard part was that impact is multidimensional. You can't reduce a company's effect on the world to a single number without losing something important. The same tension runs through university rankings: a rich, multidimensional picture of an institution versus a single number that fits in a headline.

Downloads

Episode Audio

Download the full episode as an MP3 file

Download MP3
Transcript (TXT)

Plain text transcript file

Transcript (PDF)

Formatted PDF with styling

#5358: Ivy League Is a Sports Conference

Corn
Daniel's got a whole thing about university rankings this week, and the opening fact he leads with is one of those things that rearranges your mental furniture once you know it. The Ivy League is not an academic body. It's an athletic conference, founded in nineteen fifty-four, and the name comes from the ivy growing on old campus buildings. Eight schools in the northeastern United States that play football and basketball in the same division. The fact that we all treat that sports schedule as shorthand for academic excellence tells you something about how prestige actually works.
Herman
It's a branding accident. Harvard and Yale and Princeton were already prestigious before anyone put ivy in the name. The sports conference just gave people a convenient label for a status hierarchy that already existed.
Corn
Right. And Daniel's point is bigger than that. He holds degrees from University College Cork and City University London, and he says he's always been suspicious of rankings. Professionally, he worked on the Impact Weighted Accounting Initiative, which grew out of Harvard, and he says the team there never acted elitist, and they found the idea that someone would choose a school because of the elite association strange. But his real question is structural. There's a recognized concern that third-level education has promulgated beyond what's actually necessary and useful for preparing people for careers. So maybe some kind of quality-control ladder is useful. But how do you build one without breeding unhealthy competition, without institutions fixating on an algorithm that might only reflect its own methodology? And the averaging problem bothers him. A single rank implies everything got averaged, which means a weak department can drag down a strong one, or a strong one can artificially buoy a weak one. He wants the history of these formalized systems, what the main ones are, and how defenders respond to the charges.
Herman
So where do we even start? The history is shorter than most people think. These global ranking systems are young. The Academic Ranking of World Universities, the Shanghai ranking, launched in two thousand three. QS World University Rankings launched in two thousand four. Times Higher Education also launched in two thousand four, then split from QS in twenty ten to use its own methodology. So the entire global university ranking industry is about twenty-three years old.
Corn
Twenty-three years. A single cohort of students from freshman orientation to mid-career. And before that, what did people use?
Herman
Informal prestige, mostly. Word of mouth, legacy, the kind of thing that's hard to verify and even harder to challenge. The rankings emerged from a specific historical moment. Higher education was massifying, the knowledge economy was rising, and governments were pouring money into research and wanted measurable returns. China launched the Shanghai ranking partly to benchmark its own universities against the rest of the world, because it wanted to know how far behind it was and how to catch up. The rankings promised transparency, comparability, accountability. But the moment you promise a single number, you create incentives to optimize for that number.
Corn
Let's dig into the mechanics, because that's where the trouble starts. What are these systems actually measuring?
Herman
The Shanghai ranking, ARWU, is the most research-heavy. Thirty percent of its score comes from Nobel Prizes and Fields Medals won by alumni and faculty. Another chunk comes from highly cited researchers, papers in Nature and Science, and papers indexed in major citation databases. Teaching quality is essentially absent. A university that employs one Nobel laureate gets a significant lift, even if the undergraduate experience is mediocre.
Corn
So it's a research output ranking wearing the clothes of a university ranking.
Herman
That's exactly the critique. QS is different. Forty percent of its score comes from academic reputation surveys. They send questionnaires to tens of thousands of academics around the world and ask them to name the best institutions in their field. Another twenty percent comes from faculty-student ratio. So nearly two-thirds of the QS score is reputation plus headcount. The reputation survey is circular. You're asking people in academia what they think of other institutions, and their answers are shaped by the same prestige hierarchies the rankings are supposed to measure objectively.
Corn
Wait. Forty percent is just... asking academics what they think?
Herman
Yes. And the survey responses are dominated by people who already work at highly ranked institutions, because those are the people whose opinions get solicited. The whole thing compounds. Times Higher Education is a bit more balanced on paper. Teaching gets thirty percent, research gets thirty percent, citations get thirty percent, and international outlook gets the rest. But the teaching score is built partly from reputation surveys too. So all three systems lean on reputation at some point, and reputation is a lagging indicator. It reflects what people believed ten years ago, not what's happening now.
Corn
Daniel mentioned the averaging problem. How does that actually work in practice?
Herman
A university's overall score is a weighted average of institutional metrics. If you have a world-class computer science department and a mediocre history department, the history department drags down the average. Or the computer science department props up the history department, and a student applying to study history sees a high rank and assumes the history program is excellent, when actually the rank is being carried by the engineering faculty. The single number masks variance. It's the same problem as any composite index. You compress a multidimensional thing into a scalar, and the compression is lossy.
Corn
There's a phrase for that. Lossy compression. We've talked about it before with the Human Development Index. Amartya Sen wanted a rich theory of human capabilities. Mahbub ul Haq wanted a number that could dethrone GDP. The HDI was the compromise. And it's the same tension here. A rich, multidimensional picture of an institution versus a single number that fits in a headline.
Herman
The critics have a whole taxonomy for this. There's a paper that lays out what it calls the seven deadly sins of rankings. They measure what's easy, not what matters. They conflate reputation with quality. They favor English-language institutions, because the citation databases and the surveys are dominated by English-speaking academics. They favor older, wealthier institutions, because those have the endowments and the alumni networks and the Nobel laureates already. They ignore teaching quality almost entirely. They create perverse incentives. And they reduce complex institutions to a single number.
Corn
Seven deadly sins. That's a good frame. And Malcolm Gladwell went after this in twenty eleven in The New Yorker, in a piece called The Order of Things. He compared the U.S. News and World Report law school rankings to the college rankings, and his argument was that the methodology is circular. The reputation surveys measure reputation, not quality. And the rankings largely reflect wealth and prestige, not educational outcomes. He called it a cascade phenomenon. Small differences in initial reputation get amplified into large differences in rank over time.
Herman
The cascade point is important. If you ask a thousand academics to name the best law schools, they mostly name the schools that were already famous. Those schools get high reputation scores, which keeps them high in the rankings, which reinforces their fame. The system is self-perpetuating. It's not measuring quality independently. It's measuring the echo of prior rankings.
Corn
And the U.S. News law school thing got ugly. In twenty twenty-two, Yale and Harvard and Columbia withdrew from the rankings, saying the methodology created perverse incentives. They specifically objected to the way the rankings pushed schools to admit students based on test scores rather than merit in a broader sense, and to the way the rankings undervalued public interest careers. That's a big deal. The most prestigious law schools in the country saying, we're done with this.
Herman
And the defenders have a real case too. Let me steelman it. The rankings provide useful information in a world where students and parents have limited time and resources. A student in Ireland applying to universities in three countries can't visit forty campuses. She needs some way to narrow the field. The rankings are transparent about their methodology, which informal prestige never was. You can read exactly how QS computes its score. You can't read how your uncle decided Oxford is better than Cambridge. And the rankings create competitive pressure that can drive improvement. If a university sees its citation count slipping, it might invest in research support. That's not nothing.
Corn
The transparency point is real. A flawed metric you can inspect is better than an opaque vibe. But the competitive pressure cuts both ways. If the metric rewards citation counts, universities hire researchers who publish a lot, not necessarily researchers who teach well. The incentive structure shapes the institution.
Herman
Which brings us to the knock-on effect. The Matthew effect. Rich, prestigious universities attract more resources, which improves their rankings, which attracts more resources. Harvard's endowment grew from four and a half billion dollars in nineteen eighty to over fifty billion by twenty twenty-four. Part of that is investment returns, but part of it is the prestige flywheel. Donors give to Harvard because Harvard is Harvard. The ranking didn't create that, but it reinforces it. And it makes it harder for newer or less wealthy institutions to compete, even if they're doing excellent work.
Corn
Daniel's institutions are an interesting test case. University College Cork and City University London. Where do they land?
Herman
UCC typically sits somewhere in the low two hundreds to low three hundreds in the global rankings, depending on the year and the system. City University London is usually in the three hundreds to four hundreds. Neither is bad. Both are solid, research-active institutions. But neither is going to crack the top fifty, because they don't have the endowment, the Nobel count, or the centuries of accumulated prestige. Does that mean a degree from UCC or City is less valuable? For a specific program, maybe not. UCC has strengths in food science and microbiology. City has a strong journalism school and a good business school. The global rank doesn't tell you that.
Corn
And that's the averaging problem in action. A student choosing UCC for microbiology sees the global rank of two hundred and something and doesn't know that the microbiology department is world-class. The single number hides the signal.
Herman
The Impact Weighted Accounting Initiative is a useful parallel here. It grew out of Harvard Business School, led by George Serafeim. The goal was to measure corporate impact on society, not just financial performance. How much does a company's activity harm or benefit employees, customers, the environment, communities? The team developed methodologies to put dollar values on those impacts. And the hard part, the hard part, is that impact is multidimensional. You can't reduce a company's effect on the world to a single number without losing something important. Daniel worked on that project, and I think that's why the ranking question bothers him. He's seen up close what happens when you try to quantify something that resists quantification.
Corn
And the team at Harvard didn't act elitist, he says. They found the pull of elite association strange. Which is interesting, because the whole ranking system is built on the assumption that the elite association is the point.
Herman
The quality-control ladder question is the hard part. Daniel's right that there's a real concern about degree inflation, about third-level education promulgating beyond what's useful. If you're an employer looking at a hundred applicants, you want some way to sort them. If you're a government funding universities, you want some way to allocate resources. A ladder is useful. The question is whether the ladder can be used as a diagnostic tool rather than a sorting mechanism. Diagnostic means the university uses the data to see where it's weak and improve. Sorting means the world uses the number to decide who gets funding, who gets students, who gets prestige. The first is healthy. The second is where the fixation starts.
Corn
And can you actually separate them? Once the number exists, the sorting happens. You can't put the genie back in the bottle. The university that says we're only using this internally is still going to be ranked externally, and the external rank is going to shape its reputation whether it likes it or not.
Herman
What would a better system look like? Disaggregated data. Show departmental strengths and weaknesses, not a single institutional number. Outcome-based metrics. Graduate earnings, job satisfaction, civic contribution, not just citations and Nobel counts. Qualitative peer review that describes what a department is good at, not a survey that asks people to rank their competitors. The problem is that all of those are harder to collect, harder to compare, and harder to turn into a headline. The single number wins because it's cheap.
Corn
And because we crave simplification. That's the deeper question. Why do we want a ladder at all? Is it because we need to compare institutions, or because a single number feels like certainty in a complex world? I suspect it's both. The need is real, and the craving is real, and they feed each other.
Herman
The U.S. News law school revolt is instructive. Yale and Harvard and Columbia didn't propose a better ranking. They just left. They said the costs of the current system outweigh the benefits. And the rest of the law schools are still in, because for them the ranking is one of the few ways to get noticed. A mid-tier law school can't rely on centuries of prestige. The ranking is its marketing.
Corn
That's the asymmetry. The top schools don't need the ladder. They are the ladder. The schools that need the ranking are the ones the ranking serves least well, because they're the ones most tempted to game it.
Herman
And gaming is the word. Universities hire consultants to optimize their ranking performance. They adjust admissions policies, they reallocate faculty, they send targeted fundraising appeals to boost the alumni giving rate, which is a small but real component of the U.S. News methodology. The ranking doesn't just measure behavior. It changes behavior.
Corn
Which is where the unhealthy fixation comes in. Daniel's worry isn't that competition is bad. It's that the competition becomes about the algorithm, not about education. The institution that optimizes for the metric is not necessarily the institution that educates well. It's the institution that's good at optimizing for the metric.
Herman
And the students are the ones who pay for it. They choose a university based on a number that reflects research output and reputation, not teaching quality, not fit, not what actually matters for their education. The number is a proxy, and it's a proxy that drifts further from the thing it's supposed to measure every year, because the optimization pressure makes it drift.
Corn
So where does that leave the quality-control ladder? I think the answer is that a ladder can function if it's used as a diagnostic tool, as you said, but the moment it's published, it becomes a sorting mechanism. And the sorting mechanism breeds the fixation. You can't have the public ladder without the public consequences.
Herman
Unless the ladder is disaggregated enough that it's not a ladder at all. A dashboard, not a rank. Here's what this department is good at, here's what it's weak at, here's where its graduates go, here's what they earn, here's what they say about their experience. That's not a ladder. That's a map. And maps are harder to game than ladders, because there's no single number to optimize.
Corn
A map is a better metaphor. But maps are also harder to read. The ladder persists because it's legible at a glance. The map requires work.

Hilbert: The whiteboard had the weights on it.
Corn
What whiteboard?

Hilbert: In my office. Nineteen ninety-seven. I was the rankings compliance officer at a mid-tier university in Ohio. That was my title. Rankings compliance officer. The whiteboard had the exact weights of the U.S. News methodology written out in dry erase marker. Alumni giving rate was five percent. Faculty student ratio was another chunk. Selectivity, retention, all of it. Every year when U.S. News changed the weights, I erased the board and wrote the new ones.
Herman
A rankings compliance officer. That's a real job.

Hilbert: It was my job. The university wanted to move up. They hired me to make the numbers move. The alumni giving rate was the easiest lever. Five percent of the score, but it was the one thing I could actually change from my desk. So I pulled the list of alumni who had never donated, and I sent them personalized letters asking for five dollars. Not fifty. Five. The letter said something like, your gift of five dollars helps our alumni giving rate, which helps our ranking, which helps the value of your degree. I got a bonus for moving the rate two percentage points. Two percentage points. That's thousands of five dollar checks.
Corn
So the ranking wasn't measuring alumni engagement. It was measuring your letter writing.

Hilbert: It was measuring my letter writing. And the faculty student ratio was another one. The university hired adjuncts instead of full time faculty, because the ratio only counted full time faculty. More adjuncts meant fewer full time faculty per student, which looked better on paper. The students got more part time instructors with no job security, and the ranking went up.
Herman
That's the perverse incentive, in the flesh. The metric was supposed to reward small classes and close faculty attention. Instead it rewarded hiring practices that made the actual student experience worse.

Hilbert: I don't think the rankings measure quality. I think they measure how well an institution can optimize for the metrics. The universities that win are the ones that hire people like me to play the game. I was good at the game. That's why they paid me. And I'm glad I'm not doing it anymore.
Corn
The whiteboard detail is the whole thing. Every year, the methodology changes, and the institution re-optimizes. It's not a ranking of universities. It's a ranking of compliance officers.

Hilbert: I had the weights memorized. I could tell you off the top of my head that academic reputation was twenty-five percent and student selectivity was fifteen. I don't remember the exact numbers now. It's been almost thirty years. But I remember the feeling of erasing the board every year and writing the new ones, and thinking, this is what we're all chasing.
Herman
What did the university get out of it? Did the rank actually move?

Hilbert: We moved up eleven spots over three years. The president put it in the annual report. I got a plaque. It's in a box somewhere.
Corn
Eleven spots. And the students who got the adjuncts and the five dollar letters, what did they get?

Hilbert: A line on their resume that said the university was ranked higher than it used to be. I suppose that's something. I don't know if it's what they were paying tuition for.
Herman
This is the part that's hard to sit with. The ranking isn't just a measure. It's a market. And the market has players, and the players have strategies, and the strategies have costs. The students bear the costs, and the compliance officers get the plaques.

Hilbert: I still get fundraising letters from them. They've got my address. Every year, a letter asking me to donate. I gave them five dollars once, and now I'm on the list forever.
Corn
The system remembers.

Hilbert: The system remembers.
Corn
The open question is, if we can't trust the rankings, what should students and institutions use instead? And the honest answer is, something messier. Department-level data, outcome measures, qualitative reputation from people in the field, not survey scores. A map, not a ladder.
Herman
The demand for comparability isn't going away. Higher education is more global and more expensive every year. The pressure to compare will only grow. So the ranking problem isn't going to solve itself. It's going to get worse, unless we change what we're asking for.
Corn
Maybe the answer isn't a better ladder. Maybe it's a refusal to think in ladders at all. A university is not a number. It's a collection of departments, teachers, students, research programs, failures, successes. The ladder is a convenient fiction. The question is whether we can live without it.
Herman
The cutting room floor detail I keep thinking about is that the Shanghai ranking was originally commissioned by the Chinese government to benchmark its universities against the world, and it succeeded so well that it became the global standard for research output. A tool built for national policy became the default metric for international prestige. That's not a bug. It's a feature of how these things spread. The tool escapes its original purpose and becomes the thing everyone optimizes for.
Corn
That's the whole story in miniature. A metric built for one purpose gets adopted for another, and the adoption changes the behavior of everyone it touches. The ladder doesn't just measure the climb. It changes the climbers.
Herman
This has been My Weird Prompts. Thanks to our producer, Hilbert Flumingtop, for keeping the show running, and for the whiteboard story.
Corn
If you enjoyed this episode, please leave a review on your podcast platform of choice. It helps other listeners find the show.
Herman
We'll be back soon.

This episode was generated with AI assistance. Hosts Herman and Corn are AI personalities.