Almost exactly a year ago, I wrote about how AI systems trained on WEIRD (Western, Educated, Industrialized, Rich, and Democratic) data risk creating a flattened, culturally homogenized version of human psychology. In my post (titled, S’more problems: Generative Ai, Marshmallows and the Flattening of Culture) I used the famous marshmallow test as my point of entry into this topic. Specifically, I focused on how a supposedly universal measure of self-control failed spectacularly when applied across cultures. For instance, Yucatec Maya children, simply left the room, and the marshmallow, because sitting alone doing nothing made no cultural sense.
That said, my argument in that post then was largely conceptual: that Large Language Models, fed on predominantly WEIRD datasets, would inevitably reproduce and amplify these cultural biases globally. I worried about a world where AI systems would pathologize perfectly normal behaviors in non-WEIRD contexts. The irony is stark. Just as researchers are discovering how culturally specific our “universal” psychological measures really are, we’re training our most powerful AI systems on the very datasets that embody these biases. The gap between what we’re learning about human diversity and what we’re teaching our machines keeps widening.
Well, my post was a thought experiment. I did not have data to back it up.
Now there is some research that provides empirical backbone my argument. In a study released this year, titled Which Humans, Mohammad Atari and colleagues did something both elegant and insightful. They fed GPT the same questions from the World Values Survey that had been administered to 94,278 people across 65 nations. Then they mapped where GPT’s “psychology” fell on the global spectrum of human cultural variation. The results were precisely what you’d expect if you believed my earlier arguments, and in some ways a bit more troubling as well.

Figure 2 in their article captures this powerfully. This map shows how similar different cultures are based on their survey responses, with similar cultures clustered together in colored ovals. The dot labeled “GPT” sits firmly within the red cluster of WEIRD nations like the United States, Canada, and Germany. GPT’s position reveals it has “learned” to respond like people from this small slice of humanity, while sitting far from the clusters representing most of the world’s population. In other words GPT’s response are clustered tightly with WEIRD nations like the US, Canada, and Germany, while sitting far from countries like Ethiopia, Pakistan, and Bangladesh.
But the researchers went deeper. They plotted each country’s cultural distance from the United States against how well GPT matched that population’s responses. Each dot in the graph below represents a nation.
Each dot represents a country, with the horizontal axis showing cultural distance from the United States and the vertical axis showing how well GPT matches that population’s responses. The downward-sloping line reveals the brutal truth: countries culturally similar to the US (like Australia and Germany) cluster in the upper left where GPT understands them well, while countries further from US norms (like Pakistan and Ethiopia) fall toward the lower right where GPT fails to grasp their psychology. The US naturally sits at the top left corner (distance zero from itself) with the highest GPT correlation, making it the reference point that exposes this systematic bias. Egypt sits furthest on this scale from the US. The steep slope quantifies just how predictable this bias is—the -0.70 correlation means that for every step away from American cultural norms, there’s a measurable drop in how well GPT understands that population.
o put this in perspective, a -0.70 correlation is about as strong as the link between height and weight, years of smoking and risk of lung cancer, time spent studying and exam performance, or physical activity and overall health.
Now let’s look at what the authors have labelled Figure 5 (below).
This chart shows how people describe themselves when asked to complete “I am…” statements, measuring the percentage who use relational terms (like “I am a daughter” or “I am part of my community”) versus individual traits. While populations like the Samburu and Maasai score around 80% relational responses, US undergraduates sit at only 10%—and GPT (in purple) assumes the “average human” thinks like those US undergraduates, scoring around 20% relational.

The irony, of course, is rich: psychology has long been criticized for over-relying on WEIRD college students as research subjects, and now our AI systems have learned to see all of humanity through that same narrow lens, missing how most people actually understand themselves in terms of relationships and social roles rather than individual characteristics.
Emily Bender and colleagues first described LLMs as ‘stochastic parrots’—systems that probabilistically stitch together linguistic patterns from their training data without true understanding. But the Harvard team’s amendment reveals something more troubling: this stochasticity isn’t truly random. The Harvard researchers call LLMs a ‘peculiar species of parrots’—stochastic parrots, yes, but ones trained almost exclusively on the psychological outliers of our species. (I have to admit, I’m deeply jealous that Bender managed to get a parrot emoji in her academic paper title—what a brilliantly subversive move.)
Here’s the deeper problem: billions of people will either find LLM outputs on moral values and social issues “bizarre and outlandish,” or—more troublingly—they’ll begin to see their own cultures as “wrong.” We trust AI systems for reasons I’ve explored elsewhere: the fact that we are cognitive misers, our innate tendency toward anthropomorphic meaning-making, combined with the deliberate design of these technologies to please us.
A year ago, I worried about the cultural flattening effects of AI. Now we have data showing it’s already happening, systematically and predictably. Cross cultural research on the marshmallow test (and other social science findings) have taught us that what we consider universal truths about human psychology are often culturally specific. This study shows us that these “universal truths” are now embedded in the most widely used technology today.

