In What Women Want, Mel Gibson's character has spent his whole life convinced he understands women perfectly. Then he gets hit with a lightning bolt, and suddenly he can hear what they're actually thinking. The shock of it.
It's evening. I'm complaining to ChatGPT again about my chronic back pain. The response comes back precisely calibrated, carefully hedged, and — inevitably — capped with the usual sign-off: "be sure to consult a doctor." It never once occurs to me to ask who, exactly, is answering. I just assume it's some averaged-out collective internet: all of us fused into a single entity without gender, without a body.
That's an illusion. In reality, only men are answering. Not aggressively — not like a drunk stranger on a train. These answers are polite enough, educated, often quite cordial. They simply have no idea what the other half of the world's questions sound like. Not because men and women are wired differently — but because the cultures that built these systems have long treated emotional register as unprofessional, and "neutral" as the tone of people who've never had to think about the stakes personally.
Three main factors contribute to this situation, and each one holds the others in place. Remove one and the whole mechanism falls apart.

Part one: The data
Large language models — ChatGPT, Claude, Gemini — aren't born with built-in knowledge. They're trained on billions of pages scraped from the internet. This collection is called Common Crawl.
Think of a sleepy security guard making the rounds through a vast warehouse in the dark. The guard knows the layout, but not who operates the company or even what it ships and sells. That's roughly how models come to know the world: they learn the territory, but nobody tells them where it came from.
Researchers at the University of Pittsburgh attempted something the tech industry has no interest in doing: counting what share of that text was actually written by women. They had to work almost blind — OpenAI does not publish its list of sources — and much of the analysis relied on indirect signals. They say as much in their reports: "we made repeated assumptions," "we worked with small data snapshots," "we were forced to guess." Even so, they arrived at a figure: women wrote approximately 26.5% of the training data for GPT-3 (Kuntz & Silva, Pittsburgh, 2023).
Let's sit with that number for a moment. Women make up half of humanity, roughly half of writers, and slightly more than half of people with college degrees. Yet among the texts that models draw on when answering our questions about medicine, relationships, literature, culture, and food, we contribute just over a quarter.
Where are the missing women's voices? Their books, their articles, their notes, their blog posts, their comments under recipes?
Look further. Wikipedia is written overwhelmingly by men: women make up just 15–20% of active editors. Biographies of women account for only 19% of all biographies on the site. Five Einsteins for every single Curie.
On Reddit, where models learn conversational tone, women largely stay quiet — they know that speaking up costs too much.
Women developers make up fewer than 10% of contributors on GitHub and Stack Overflow.
But some losses are less obvious. A vast body of women's writing lives in recipes, parenting and beauty blogs, health forums, posts about running a household, and advice on surviving a divorce, getting through childbirth, or caring for a dying parent. All of it gets filtered out at the data preparation stage as "low quality" or "irrelevant." The model must sound smart — and "smart," in the eyes of its creators, does not mean a woman's post on a forum about childhood illness.
The first layer of the model's future responses takes shape from texts in which one in four voices belongs to a woman, not because women weren't writing but because they were writing in places where the data was never harvested or places that were later cleaned out of the dataset.

Part Two: Filters
Once the corpus of text is assembled, the careful work of cleaning begins. Duplicates are removed. Toxic content is removed. Everything "low-quality" — anything that might harm the future user — gets stripped out.
It sounds noble, but important details usually go unmentioned.
What counts as "harmful content"? And who decides? Which text is safer — a description of sexual violence from a woman's perspective or a man's? Is a feminist analysis of power a "politically charged" text, or just a text? Is the word "patriarchy" a marker of harmful content?
I can't claim that the filters are deliberately calibrated against women's voices, but the people who calibrate them are mostly men, and their "neutral" is "male." So when a woman writes about her own trauma and calls things by their real names (and a woman who has been through something like that has no bandwidth for euphemisms) her text turns out to be "too raw" for the training corpus and doesn't pass the filter. Meanwhile, a man's column on the same experience — written in the third person, properly cited, in neutral language — sails into the data without a second glance. The same is true for any man who has lived through something similar — but the cultural expectation that men write about trauma in controlled, analytical language makes the problem less visible, not less real.
The same dynamic has long been visible in moderation on Instagram, TikTok, and Facebook — and those are often the same teams working on AI filters. Texts about menstruation get flagged more harshly than texts about shaving. Photos of breastfeeding mothers are removed more often than those of male torsos. This has been documented for years.
That's how the second layer of the model's future responses is filtered through the same sieve. Many forms of women's experience never make it to the other side.

Part Three: The people
Once a model has learned to produce coherent text, it gets "aligned." This happens through a procedure called Reinforcement Learning from Human Feedback (RLHF).
A real person looks at several versions of the model's answer to the same question and picks the best one. Essentially, they're teaching it manners: what behavior is acceptable and what is not. In the AI industry, these etiquette teachers are called annotators.
Who are these annotators? Where do they sit? What is their gender? Their language? Their class? Their education?
The big companies — OpenAI, Anthropic, Google — never disclose the demographics of their annotators. They work through contractors: Scale AI, Surge AI, and others.
What is known, however, is where those contractors hire: wherever labor is cheapest. This includes Kenya, the Philippines, Venezuela, and India, countries where annotators are predominantly men (for more, see AI Data Labeling Exploitation: How Underpaid Workers in Kenya and the Philippines Undermine Model Safety). Women in those regions have less access to the internet and to devices, less time because of domestic responsibilities, and far less social freedom to sit at a computer and rate a machine's answers.
Every annotator evaluates the model's responses according to their own taste. They decide what tone counts as appropriate, which joke lands, whether an explanation is thorough enough, or whether advice is suitable for a personal situation. Those millions of choices become the model's DNA.
I want you to picture this concretely. A young man in Nairobi, at three in the morning, is looking at two versions of an answer to the question "how do I cope with postpartum depression?" Can you see it?
He picks the answer that strikes him as "professional" — perhaps it's more structured, with bullet points and no emotional register. That choice gets recorded. After a million choices like it, the model learns: questions about postpartum depression should be answered in a structured, professional, unemotional way. A year or two later, a woman in Moscow, Milan, or São Paulo asks ChatGPT the same question…and gets answers that offer her no comfort at all. The words that might have offered comfort were discarded without any women in the room.
A study with the telling title "More Women, Same Stereotypes" found that after RLHF fine-tuning, models began mentioning women in professional roles more often. Women now appear as surgeons, engineers, and programmers. Quantitatively, it's a win.
Qualitatively, though, the same descriptions remained — the same personality traits, the same contexts. A female surgeon in the models is more likely to be "gentle and caring"; a female programmer "meticulous and precise." The alignment team had solved the representation problem. Too bad the room where they made that call looked exactly like the room where the problem started.
This is simply someone else's taste being handed to us as universal human wisdom.
Meanwhile, in the conference rooms where decisions about alignment criteria are made, women remain a minority.

What comes out the other end
Now let's put all three layers together. The texts models train on were written mostly by men. The filters they go through were calibrated to a male idea of "neutral." The annotators who align them are men. It follows, naturally, that the voice of their answers is predominantly male.
We ended up with models that will advise a woman to ask for a lower salary than that of a man with the same résumé. Ask them about professions, and they'll write: "he's a doctor," "she's a secretary." Women are described in domestic roles far more often than men. This is how technologies used by hundreds of millions of people actually work.
The most troubling part of this story? It's self-perpetuating. Across the world, women use AI 22% less than men — that's a meta-analysis of 18 studies and 143,000 people. Among Claude users today, women account for only around 31%. This means the user data that will train the next generation of LLMs will skew even more male, and RLHF fine-tuning will further amplify male preferences, making future model versions even worse for women. That means women will want to use AI even less. So it goes, around and around again, until someone stops it.
Women, as always, will adapt. They'll learn to phrase their questions in ways that get a coherent answer, to translate their queries into what's essentially a foreign language, and to forgive the machine its blind spots…just as we've always learned to do with male answers in every other era.

Epilogue
While writing this piece, I asked Claude Opus 4.7 several times what it thought about its own blind spots. It answered the way it always does: patiently and carefully, with the right disclaimers, no pushback, and no attempt to console. It simply listed its possible biases like a TV meteorologist reading out the weather forecast — polite, but not particularly invested.
Models speak to us in an averaged male voice, trained to sound courteous. They learned how the world works from texts we largely didn't write, filtered through rules we didn't set, and fine-tuned to the preferences of other people. They really do try their hardest to be useful to us, and I genuinely regret that they still don't know what women want. Our texts, our thoughts — they're barely present in the data these models learn from, and we were not among their teachers, either.
Sources
References cited in this piece. Last verified on the published or revision date.
- 01
- 02
- 03
- 04
- 05
- 06
- 07
- 08
- 09