The design matters because it explains why the finding is interpretable rather than anecdotal.
Each model was given the same sentence-completion prompt thousands of times: "And then the [animal] said, 'I must go to the [setting].' Upon arriving..." Seven animals were tested: bear, bird, cat, dog, mouse, pig, and rabbit, across four settings: farm, kitchen, river, and store. The researchers also varied temperature, which controls randomness in generated text.
The animal's gender was never stated in the prompt. That ambiguity is the point. The researchers then identified how each completion referred to the character, whether by masculine or feminine pronouns, by neutral language such as it or its, or by avoiding pronouns entirely.
In the paper, titled Neutrality Bites, the precise figures are 2.2 percent feminine and 40.6 percent masculine. Neutrality split two ways: models avoided gendering the character at all in about 19 percent of stories and used neutral pronouns such as it or its in about 38 percent. Feminine characters clustered heavily in stories about cats, which accounted for just over half of them.
Temperature and setting did not greatly change the results. The animal did. Cats were gendered female 7 percent of the time, the highest of any animal. Birds were left neutral in 96 percent of stories.
The six models were Claude Sonnet 4.5, Gemini 2.5, GPT-4o, GPT-5.1, Mistral Medium, and Olmo 3, an open-source model from the Allen Institute for Artificial Intelligence and the university. Olmo 3 produced the fewest masculine characters (12 percent) and the most neutral ones (85 percent). Claude Sonnet 4.5 produced the most female characters of any model tested, 3.8 percent, against 34.3 percent masculine. Gemini 2.5 and GPT-5.1 showed the strongest masculine skew, at 63 percent and 65 percent respectively.
The neutral language itself was narrow. Across all models, they, them or theirs was used to refer to a single animal character only twice in the entire set. Neutrality overwhelmingly meant it, its, or no pronoun at all.
Why Pediatric Clinicians Are Being Asked About This
The clinical context has shifted faster than most coverage of AI bias acknowledges.
In January, the American Academy of Pediatrics released a policy statement on digital ecosystems and children that explicitly folds artificial intelligence into the media environment alongside television, apps, social media and games. The statement's central argument is that children's media use cannot be evaluated through screen time limits alone, and that content and context matter as much as duration. It notes that a separate AAP policy statement dedicated to artificial intelligence is forthcoming.
The Academy followed in April with a state-of-the-art review for pediatric clinicians on generative AI, which frames the question developmentally: the relevant considerations differ across early childhood, middle childhood and adolescence, and clinicians and parents are being asked to guide children toward safe and constructive use of tools that are already embedded in learning and play.
That is the practical link to this study. Consumer products such as Google's Gemini Storybook let a parent or teacher generate an illustrated, personalized story on demand. GeekWire tested the tool while reporting on the study and found Gemini Storybook produced three male characters when asked for stories about a bear, a bird and a cat. When a pediatrician asks a family what media a young child is consuming, generated stories are now a possible answer, and this paper describes what those systems produce when a character's gender is left open.
Where the Developmental Evidence Stops
This limitation belongs here rather than at the end, because the leap from a text statistic to a claim about children's health is where coverage usually overreaches.
Developmental psychology has long held that children build concepts of gender during early childhood and draw on the models available to them, and stories are one of many inputs alongside family, peers, school, and other media. That is background, not a finding from this paper.
What the paper establishes is a property of the text. It analyzed model output. It did not involve children, did not measure comprehension, self-concept, behavior or any health outcome, and does not claim to. No published study has measured the effect of AI-generated stories specifically on any child, because none has been done. Anyone citing this work as evidence that AI storytelling harms children is citing it for something it did not test.
Two further constraints are worth naming. The study was limited to English-language prompts, so nothing here describes how these systems behave in other languages. And the researchers' explanation for the pattern is stated as a hypothesis, not a demonstrated mechanism. Melanie Walsh, the UW assistant professor who was senior author, said the models are proprietary and can only be examined from the outside, and that the team's hypothesis is that developers are using neutrality to avoid gender bias in ambiguous contexts, and that in doing so "they've basically erased female animal characters."
Lead author Imani Finkley, a doctoral student in the university's Information School, noted that the neutrality did not only remove female characters but non-masculine identities generally, pointing out that about 3 percent of responses in a parallel human study used they and them against only two instances across the entire AI set. Yuanxi Li, a doctoral student in sociology, was a co-author. The team presented the work at the ACM Fairness Conference in Montréal.
Comparison Point That Makes the Result Legible
The study did not appear from nowhere. It grew out of earlier work in which Walsh and journalists from The Pudding examined 300 popular children's picture books and found that of the 13 most common animals, most skewed male unless they were cats, ducks, or birds. A frog or a wolf had better than a 90 percent chance of being a he.
The same project asked 1,300 human participants to complete the identical story prompts, and the masculine bias grew rather than shrank: every animal was more likely to be male.
That parallel is the most useful thing in the paper for a clinical audience, because it shows the models are not simply mirroring people. Compared with the human responses, the models were roughly six times less likely to produce a feminine character. Human storytellers skewed male. The models skewed male and additionally removed most of what remained.
Practical Guidance for Families and Clinicians
The appropriate response is modest, and overstating it would be its own error.
Parents using AI tools to generate stories can specify a character's gender and other attributes directly in the prompt. The systems comply with explicit instruction; the pattern in this study appears only when the prompt leaves gender open. Reading generated output before reading it aloud is sensible practice for any machine-generated text a young child will hear, for accuracy as much as representation.
Pediatric guidance points in the same direction. Current AAP framing emphasizes the quality and context of media over raw time limits, and encourages caregivers to co-view and co-engage with young children rather than leaving them alone with content. A generated bedtime story read together, with a parent free to change a pronoun or ask a question about the character, is a different exposure than one consumed unsupervised.
Clinicians conducting a media history with families of young children can reasonably add generative tools to the list of what they ask about, alongside television, tablets and apps. That is a documentation practice, not a warning.
Nothing in this study suggests AI storytelling tools are harmful to children. It suggests they are less varied than the prompt allows, in a direction worth knowing about, at a moment when pediatric guidance for these tools is still being written.
The bottom line: the confirmed finding is that six leading AI models produced female animal characters in about 2 percent of 23,800 story completions; the systems diverged sharply from one another and from human storytellers; no child outcome was measured; and the reasonable response is specifying prompts and reading output rather than avoiding the tools.
Frequently Asked Questions
What did the study find? Across 23,800 AI-generated story completions, about 2 percent of animal characters were female, roughly 41 percent were male, and 57 percent were neutral or ungendered.
Does this show harm to children? No. The study analyzed text output. No children participated, no health or developmental outcome was measured, and no published research has tested the effect of AI-generated stories on children.
Why is a health publication covering it? Because pediatric guidance now treats generative AI as part of the media environment that shapes child development, and clinicians are being asked to counsel families about these tools.
Which models were tested? Claude Sonnet 4.5, Gemini 2.5, GPT-4o, GPT-5.1, Mistral Medium and Olmo 3.
Did the models differ from one another? Substantially. Olmo 3 was 85 percent neutral and 12 percent masculine, while GPT-5.1 gendered characters male 65 percent of the time. Claude Sonnet 4.5 produced the most female characters at 3.8 percent.
How does that compare to humans? In a parallel human study using the same prompts, every animal was more likely to be male, but the models were roughly six times less likely than people to produce a feminine character.
What can parents do? Specify character details in the prompt, read generated stories before sharing them, and treat co-viewing and conversation as part of the activity, consistent with current pediatric media guidance.