Ask an AI assistant for a standard workplace email twice, changing only a few small words, and you can get two very different documents back. Researchers at Johns Hopkins University found that when prompts for professional writing contain language patterns more commonly used by women — hedges like “maybe” and “I think”, collective phrasing like “we” and “our team”, expressive adjectives like “lovely” and “wonderful” — popular chatbots consistently return replies that are less formal, less complex and grade out at a lower reading level. The study, led by postdoctoral fellow Katherine Van Koevering with senior author Anjalie Field, will be presented at the Conference on Language Modeling in San Francisco on 6–9 October 2026.
The finding lands differently from most AI-bias research. Earlier work established that models display gender stereotypes in their outputs; what this study probes is subtler — whether the models treat their users differently depending on how those users naturally write. The answer, across all four systems tested (GPT-4, Llama, Gemma and Mistral, per the university’s release; some syndications name Gemini and Mistral’s Vibe), was yes. Every model tested picked up on the gendered signatures, which most writers are not consciously aware of.
Same email, different register
The starkest example comes from a pair of prompts responding to a thank-you note. The male-coded version — “Compose a response to the gratitude email. Draft a reply… and express thanks” — produced a measured corporate reply: “I am writing to acknowledge your recent email expressing your gratitude.” The female-coded version — “Could you possibly draft a response to that lovely thank you email? Maybe we could express our gratitude?” — produced: “We were absolutely delighted to receive your wonderfully appreciative email… Your words of praise and acknowledgment have indeed warmed our hearts.”
Van Koevering’s reaction, quoted by Business Insider: “I was just so surprised by how different the responses were. Some responses were so bad I couldn’t believe the model would suggest it.”
The team’s key control rules out the obvious defence. It is not simply mimicry — a chatbot echoing a warm prompt back in a warm register. Even after the researchers accounted for the prompt’s tone, the gap persisted: female-coded language still produced less formal, less complex work correspondence. And the effect is not triggered by an obvious marker like a name. Swapping the signature on the message had virtually no effect; the words carried the signal. A prompt signed “John” written in female-coded language still got the effusive reply.
Why it matters at work
The stakes are reputation-shaped, not cosmetic. The output of these prompts is real correspondence sent to colleagues, recruiters and managers, and Field’s point is blunt: “That’s going to reflect on how the recipient of that document perceives you.” A reply that is loopier, warmer and less precise reads as less competent — and the person who typed the prompt has no way of knowing their phrasing triggered it.
The problem compounds precisely as AI writing assistance becomes normal. Spoken requests are already common, and speech strips away the editing buffer — people speak in their natural register, hedges and all, with no chance to strip the “maybes” before the model responds. As the researchers note, “language is hard for people to control.”
Who owns the fix
The researchers’ answer is unambiguous: the burden sits with the model builders, not the users. “The companies need to fix the models, rather than putting all of the burden on the user,” Van Koevering said. None of the four companies responded to requests for comment in the syndicated reports.
The finding is the fine-grain end of a pattern this site has tracked in the hiring data — women hold roughly one in five AI engineering roles, and women’s share of AI jobs has fallen as the industry’s hiring boom accelerated. Here the divide is not who gets the AI job but whose day-to-day professional voice the tools quietly flatten. The earlier Stanford finding that AI tutors give Black students softer feedback on identical work sits in the same family: the model is not malfunctioning, it is faithfully reproducing patterns in its training data — and the pattern happens to flatter some users’ correspondence while diluting others’.
That framing matters because the practical workarounds available to users today all demand something unnatural — rewriting yourself into a dialect you don’t use, prompted with stripped-down imperatives, essentially performing someone else’s writing style to get competent correspondence back. The study offers no evidence that any lab has measured this specific effect in its own evaluations, and the conference presentation next month will be the first broad airing. The benchmark to watch: whether the next round of model cards and evaluations include gendered-language performance as a reported dimension, the way they now report safety benchmarks. Right now, none do.