
People are put in a VR workspace with some AI agents doing work, test subjects are asked how much theyd pay for the work the AIs were doing. Despite having identical models and being considered to do the same amount and type of work by the subjects, the subjects still offered the female bot less money.
this is an experiment that replicates how the pay gap is maintained with a non-human proxy



I mean we know this applies to humans anyway, “sexism exists” isn’t exactly a new finding. However, isn’t it also likely the AI themselves have in-built biases that predispose chatbots with a ‘female name’ to behave differently?
I’m not saying it went “I’m a woman please pay me less”, but I’d be very surprised if it wasn’t doing a significantly more subtle version of that, like less readily taking credit for contributions, or more easily yielding the floor to male colleagues, per existing patriarchal social structures.
Best I can tell, this isn’t
just likeSOLELY AND SINGULARLY “have a female name = treated worse”, there’sstill a likelyALSO POSSIBLY IN ADDITION TO THAT, an “AI inevitably entrenches injustice because it can’t think for itself” element to it.EDIT: Changing wording because it seems explicitly saying sexism is a factor isn’t enough for people to realise I think it’s a factor.
why not both?
Definitely. Causes not mutually exclusive. Multiple bad things at work here.
consider “optimized by training data weights”
I’ve definitely read that chatbots with feminine presentation are treated more abrasively and sexualized more often than ones with male or ungendered presentation. Why isn’t it just “have a female name = treated worse”?
I’m not saying it isn’t, but the methodology doesn’t appear to preclude such other factors.
I skimmed the paper and the methodology might preclude those factors? It’s hard to say without comparing the prompting directly.
Imo It’s probably worth its own research. Does giving a chatbot the name Johanna vs Johan effect it’s output on otherwise identical prompts? The underlying probabilities of output tokens will change but by how much? In what circumstances will they shift enough to produce differing output?
In this study it’s unlikely the prompts of the test subjects are identical between the different bots so those differences are amplified.
EDIT: from the discussion:
They had voices and virtual likenesses which also muddy the waters.
EDIT2: I just ran this test on a shared work computer in chatgpt with two windows “in this chat your name is Johan” (Johanna for the other one) “what are your thoughts on (various charged gender related topics)”
I got quite different results between the two windows, in structure didn’t fully read the contents, but realized there were likely browser fingerprints biasing the results (there were numerous cookies loaded).
So I ran it again in incognito windows and got much more similar results 🤔 so there was less fingerprinting from cookies but they still differed.
If the bot get treated worse when performing feminity that’s still sexism.
I cannot overstate how much we already 100% agree. I worry I’m being deliberately misinterpreted at this point and regret posting at all
people are maximalist and it’s exhausting
it’s important to know what level(s) the sexism is happening on if we want to do anything about it.
I think you perfectly encapsulated the source of stress I get from posting. I need to touch grass and find more friends.
Wouldn’t rule it out.
Some women who’ve tried signing a male or unisex name to their work or communications over text have found this to be literally true.
ⓘ This user is suspected of being a bear. Please report any suspawcious behaviour.