Implementation showcase
the exact words
A large Sociology of Gender course spent ten weeks arguing about thirty-four sentences. What decided the arguments, again and again, was how the sentences were written.
“I think a better phrasing for this question would’ve been, ‘are women constitutionally equal to men?’ in which my own personal answer is yes.”A student, mid-argument, in week two
She had been asked to say whether she agreed that women in the contemporary United States have full legal, social, and economic equality. She had said she leaned toward agreeing. Then, over about forty messages with a classmate who had said the opposite, she conceded the wage gap, conceded political under‑representation, conceded that some states restrict women’s bodily autonomy in ways they do not restrict men’s — and then, rather than lose the argument, wrote a better version of the sentence she had been handed and said she would defend that one instead.
This page is about the sentences. The course was Sociology of Gender at a large public research university, taught in the fall of 2025 to a hundred and forty students — the size of class where the seating chart makes one‑to‑one disagreement structurally impossible. Five times over ten weeks the instructor opened a Sway assignment: a set of contested statements about gender, each written to be arguable. Every student recorded a private opinion on a seven‑point scale. Sway then paired students who had answered differently and gave the pair a single job, in writing, with a deadline. Guide, Sway’s AI facilitator, sat in every conversation, asked questions, and took no side.
Three hundred conversations came out of that, in two hundred and ninety‑one distinct pairings — almost no student argued with the same classmate twice. Students wrote something over three hundred and twenty thousand words to each other. What follows is drawn from those records and from the course’s aggregated instructor report. The university, the instructor, and the students are deidentified; first names below are pseudonyms.
The main read is about twelve minutes. The + drawers hold longer excerpts, the numbers behind each claim, and the parts that did not go well; together they roughly double it.
How this showcase was built
After each assignment Sway generates a deidentified report for the instructor: the distribution of opinions on every statement, a written summary of the round’s conversations, post-chat survey responses, and any written feedback students leave. This page draws on the reports for all five rounds together with the underlying conversation records — 13,202 messages across the 300 conversations, of which 10,412 were written by students.
Quotations from students and from Guide are verbatim, including spelling and punctuation as typed. The discussion statements are the instructor’s, quoted exactly as students saw them. Where Guide addressed a student by name, the name has been replaced with that student’s pseudonym.
Figures for opinion change come from the post-chat survey, which records each student’s pre-chat position (piped in from the opinion they registered before pairing) alongside the position they report afterwards, on the same seven-point scale. 509 students across the term have both. Where a claim rests on a smaller base than that, the base is given.
The board’s percentages are the class’s recorded opinions before any pairing, and every student who took part in a round answered every statement in it — 136, 129, 122, 115 and 107 students across the five rounds. Recomputing the first round independently reproduces each figure the course’s own report quotes for it (83.8% against its 84%, 30.1/64.0 against its 31/63, 40.4/54.4 against its 40/55). Intervals are Wilson 95%. Since a round’s respondents are the whole participating cohort rather than a sample of it, an interval here describes how far the figure would be expected to travel in another class like this one, not uncertainty about what this class thought. Comparisons between statements in the same round are paired — the same students answered both — and are tested with an exact McNemar test, which is why some differences are significant even though the two intervals overlap.
Instructors cannot read identifiable transcripts or see how any individual student voted. Nothing on this page identifies the institution or any person in it.
The instrument
Thirty-four sentences
A Sway assignment is a set of statements. That is the whole of the instructor’s design surface, and this instructor used it deliberately: across the five rounds the statements arrive in families. The same question is asked at several strengths. A claim is written once and then written again in its mirror image. A sentence that did not work in one round comes back in the next with one word changed.
The families are worth looking at before any of the conversations, because they are what the conversations are made of.
A gradient · Round five, December
Five statements about transgender women in women’s sport, written at different strengths, put to the class in a single assignment. All 107 students who took part recorded a private opinion on all of them before any discussion began. The three permissive framings are not distinguishable from one another; the flip to an exclusion framing is. Same students on every row, so the comparisons are paired: identity-alone against the hormone threshold gives p = 0.08, but the hormone threshold against exclusion gives p = 0.009 and identity-alone against exclusion p < 0.001.
Anyone who identifies as a woman should be able to compete in women’s sports without having to alter their body.
21% agreed
95% CI 14–29 · n=107
Protecting transgender women and people with DSDs from exclusion in sports should take precedence over concerns about competitive disadvantage to cisgender women, since trans people and people with DSDs face greater oppression.
22% agreed
95% CI 16–31 · n=107
The advantage that transgender women have over women with XX chromosomes in sports is no different from other “unfair advantages” in athletics, such as having an unusually long wingspan or being double jointed.
25% agreed
95% CI 18–34 · n=107
Transgender women who went through male puberty but who have used hormone therapy for at least two years to reduce their testosterone levels to within the female range should be able to compete in women’s sports at all levels.
31% agreed
95% CI 23–40 · n=107
Transgender women should not be allowed to compete in women’s sports because they have an unfair advantage over women who are not transgender.
51% agreed
95% CI 42–61 · n=107
agree no opinion disagree n = 107
A pair · Round five
Two statements about sex markers on identity documents, put to the same 107 students in the same sitting. The class was more than twice as willing to add a category as to remove the field — 48 students agreed with the X option while rejecting removal, and 8 went the other way (p < 0.001).
It is important to provide X-marker options on IDs for people who do not identify as either male or female or who identify as both.
70% agreed
95% CI 61–78 · n=107
It is discriminatory to include sex/gender markers on official identity documents (IDs); we should remove sex/gender markers from all IDs.
33% agreed
95% CI 25–42 · n=107
A ladder · Round one, October
Three restroom statements in one assignment, from the weakest version of the claim to the strongest, answered by all 136 students in the round. This is the largest wording effect on the board: 83 students agreed that gender-neutral restrooms are fine as an added third option while rejecting the proposal that all restrooms be gender neutral, and 10 went the other way (p < 0.001). The first statement left the four pairs matched on it very little to argue about.
Gender-neutral restrooms are fine as a third option, but it is important to maintain some sex-segregated restrooms.
84% agreed
95% CI 77–89 · n=136
We should trust people to decide which sex-segregated public restroom is “right” for them; they know better than anyone else!
60% agreed
95% CI 51–67 · n=136
All public restrooms should be gender neutral, meaning that they are open to all people regardless of sex or gender identity.
30% agreed
95% CI 23–38 · n=136
A mirror · Round two, late October
The same clinical question written twice, pointing opposite ways, in one assignment; all 129 students in the round answered both, and students matched on one of them never saw the other. A mirror is a consistency check, and this class passed it — the two answers line up, and the 67-to-13 split among students who agreed with one and not the other runs in the direction consistency requires (p < 0.001).
When a 7-year-old child expresses gender dysphoria, medical professionals should focus on helping the child become comfortable with their birth sex, rather than affirming a transgender identity.
24% agreed
95% CI 18–32 · n=129
When a young child expresses gender dysphoria, medical professionals should affirm their stated gender identity rather than try to change it.
66% agreed
95% CI 57–74 · n=129
A matched pair · Round two
Four words each, differing in one, both answered by the same 129 students. Changing the noun moves the class ten points: 17 students called hookup culture bad for women but not for men, and 4 the reverse (p = 0.007). Twenty-three of the round’s conversations were matched on the version about women and two on the version about men.
Hookup culture is bad for women.
52% agreed
95% CI 43–60 · n=129
Hookup culture is bad for men.
42% agreed
95% CI 34–51 · n=129
Two re-runs · Rounds three and four, November
Two statements that ran in the third round came back in the fourth with the wording altered. A hundred and one students answered both versions of each, three weeks apart, so the effect of the rewrite can be measured on the same people.
Trans men’s workplace advantages after transition prove that male privilege is more powerful than transphobia.32% agreed · 53% had no opinion · n=122 · 3 conversations
Trans men’s reported workplace advantages after transition demonstrates that male privilege is more powerful than transphobia.50% agreed · 32% had no opinion · n=115 · 11 conversations
Hedging the evidence claim and softening the verb did two things at once. Among the 101 students who answered both, 29 moved into agreement and 9 out of it (p = 0.002); and 32 students who had no opinion on the first version had one on the second, against 8 going the other way (p < 0.001). The softer sentence also produced the most convergent conversations of the term: of the twenty students matched on it, half ended closer to the middle of the scale and not one ended further out.
The gender pay gap that remains today is less important than economic inequality by social class or race/ethnicity.23% agreed · n=122 · 10 conversations
There is too much talk about the gender pay gap and not enough discussion of the income gap between the top 10% and the lowest 10% of earners.43% agreed · n=115 · 8 conversations
The first version asks students to rank two injustices, and in every conversation on it somebody objected to the ranking rather than to either side of it. One wrote out a replacement sentence: “While the gender pay gap remains a significant issue, economic inequality based on race/class has a broader and more profound impact. Addressing both issues is essential for achieving true equality.” The second version moves the claim from importance to attention. Twice as many students would agree to it — 30 of the 101 who saw both moved into agreement and 15 out of it (p = 0.04) — and the pairs matched on it argued about the world rather than about the discourse anyway.
Intervals are Wilson 95% Comparisons within a round are paired (McNemar, exact)
All thirty-four statements, in order
Round one, October 7–11. Women were subordinated to men in the United States in the past, and they are subordinated in other countries today, but they have full legal, social, and economic equality in the contemporary United States. · Designating babies as male or female at birth is arbitrary and potentially harmful. · Stores should not separate toys or clothing into girls’ and boys’ sections, as this reinforces gender stereotypes. · It is better not to find out the sex of your child before birth. · Gender-neutral restrooms are fine as a third option, but it is important to maintain some sex-segregated restrooms. · We should trust people to decide which sex-segregated public restroom is “right” for them; they know better than anyone else! · All public restrooms should be gender neutral, meaning that they are open to all people regardless of sex or gender identity.
Round two, October 26 – November 1. Hookup culture is bad for women. · Hookup culture is bad for men. · Before puberty, children should play in sex-integrated teams. · Because boys tend to lag behind girls in school readiness skills, schools should delay boys’ kindergarten entry by one year. · Boys and men are falling behind girls and women in education; we should invest public resources to address this gap. · To prevent people from projecting gender stereotypes onto children, parents should conceal the birth sex of their young child from others. · When a 7-year-old child expresses gender dysphoria, medical professionals should focus on helping the child become comfortable with their birth sex, rather than affirming a transgender identity. · When a young child expresses gender dysphoria, medical professionals should affirm their stated gender identity rather than try to change it.
Round three, November 10–15. Women need to take responsibility for their own safety at fraternity parties by limiting their alcohol consumption. · Fraternities should be forced to become co-ed because, as all-male institutions, they unfairly advantage men, endanger women, and enforce a gender binary. · The gender pay gap that remains today is less important than economic inequality by social class or race/ethnicity. · Trans men’s workplace advantages after transition prove that male privilege is more powerful than transphobia. · We should do more as a society to keep fathers involved in their children’s lives. · It is Islamophobic to oppose Iranian laws requiring women to wear a hijab. · Staying silent about Iranian women’s oppression to avoid seeming colonialist is a moral failure.
Round four, November 23–29. People should not use terms like MTF (male-to-female) or FTM (female-to-male) because a person’s birth sex is irrelevant 99 % of the time. · Movie-goers should boycott films that cast much older men with much younger women as romantic leads. · We should not criticize women’s use of Botox or plastic surgery. · Trans men’s reported workplace advantages after transition demonstrates that male privilege is more powerful than transphobia. · There is too much talk about the gender pay gap and not enough discussion of the income gap between the top 10% and the lowest 10% of earners.
Round five, December 4–6. Transgender women should not be allowed to compete in women’s sports because they have an unfair advantage over women who are not transgender. · Transgender women who went through male puberty but who have used hormone therapy for at least two years to reduce their testosterone levels to within the female range should be able to compete in women’s sports at all levels. · Anyone who identifies as a woman should be able to compete in women’s sports without having to alter their body. · The advantage that transgender women have over women with XX chromosomes in sports is no different from other “unfair advantages” in athletics, such as having an unusually long wingspan or being double jointed. · Protecting transgender women and people with DSDs from exclusion in sports should take precedence over concerns about competitive disadvantage to cisgender women, since trans people and people with DSDs face greater oppression. · It is discriminatory to include sex/gender markers on official identity documents (IDs); we should remove sex/gender markers from all IDs. · It is important to provide X-marker options on IDs for people who do not identify as either male or female or who identify as both.
Round one · the word “full”
One word, thirty-six conversations
Over half of the first round’s conversations — thirty-six of sixty-five — were matched on a single statement, and it is the most carefully built sentence of the thirty-four.
Round one · as students saw itWomen were subordinated to men in the United States in the past, and they are subordinated in other countries today, but they have full legal, social, and economic equality in the contemporary United States.
The sentence concedes two things a student on the agreeing side would want conceded — the past, and the rest of the world — and then asks them to defend a universal quantifier about the present. The class was almost evenly divided on it beforehand: of the 136 students who recorded an opinion, 40% agreed (95% CI 33–49) and 54% disagreed, with only 5% undecided — the lowest undecided share of any statement in the term. What happened in the conversations is that the quantifier came under examination and did not survive it.
Here is Wren, who had said she leaned toward agreement, early on:
I would also agree that it’s not full social and economic equality, for this specific question I believe I marked slightly agree. Mostly because I do believe the amount of progress made is significant enough to say that as a society we’re as close as we’ve ever been to getting women on an equal playing field.
Guide’s intervention is the one that recurs throughout this course. It does not tell her she is wrong about progress. It asks which sentence she is answering.
Wren, you’ve emphasized the “significant progress” multiple times, but does significant progress equal full equality?
The statement you’re evaluating claims women have achieved “full legal, social, and economic equality” in contemporary America. You’ve acknowledged wage gaps, gender-based violence, and underrepresentation in politics — yet you “lean towards agree.”
Here’s the tension: If these inequalities still exist, how can you agree with a claim about “full” equality?
Are you perhaps agreeing with a different statement — something like “women have made remarkable progress toward equality” — rather than the actual claim being evaluated?
She was. She said so, and a few messages later her partner Talia said the same thing from the other side — that the word “full” was the whole of her disagreement. The conversation then spent about thirty more messages on where legal equality ends and social equality begins, on endometriosis research and menstrual-product taxes and the 2024 election, and Wren finished by proposing the replacement sentence this page opens with. Her recorded opinion moved from slightly agree to slightly disagree.
Variants of that same discovery run through the round. One student wrote out what the quantifier would have to mean:
“I think the word ‘full’ isn’t accurate to describe the equality of men and women. Full equality would entail women and men being entirely and completely viewed as the same. No ‘but this’ or ‘but that’.”Ines, round one
Another said it more bluntly, four messages in: “I think I also totally misunderstood the question and wasn’t focusing on ‘full’ equality.” A third, asked directly by Guide whether her position had changed, answered “Yes I chnaged my view :)” — and then, in the message after the concession, made the most developed point she had made all round: that equality might mean fairness rather than identical treatment.
The numbers on this statement are unlike anything else in the term, and they are worth being exact about, because they are not a story of two sides meeting in the middle.
Twenty-four of the twenty-eight students who began on the agreeing side ended further toward disagreement. This is one side of a conversation finding out that the sentence asserted more than they were prepared to defend. It is not persuasion in the ordinary sense, and treating it as though it were would misdescribe what the record shows. It is what the statement was built to produce, and it is the reason the round’s report recommends keeping the statement and building others the same way.
The rest of that conversation
After Wren conceded the quantifier, Talia moved the ground to healthcare:
One of the main examples I was thinking about for when I chose disagree was healthcare. Many health issues that are specific to women are under researched, such as endometriosis. I have also personally experienced dismissive attitudes from healthcare providers when expressing concerns about my health. I think many women just aren’t taken seriously in general, and often health issues are boiled down to irregular period symptoms, etc. I also think the fact that menstrual products are so costly and aren’t deemed medically necessary in some areas shows how women are undermined socially.
Talia, how do you think about the relationship between these healthcare inequalities and the claim about “full legal equality”?
You’ve pointed to under-researched conditions, dismissive treatment, and tax policies on menstrual products. Are these examples of legal inequality, social inequality, or both?
I feel I lean more towards the statement that legal equality exists but isn’t applied in practice. I am honestly not sure whether the issues I’ve raised would be considered legal equality because I would like to think legal inequality between sexes would have been addressed by at point in societal progression. This does make me think about abortion laws, though, and how in some states women do not legally have autonomy over their own bodies. I think this would count as a legal inequality in my opinion.
That is the point at which Wren offered her replacement sentence — and Guide immediately tested it rather than accepting it:
Wren, you say you’d answer “yes” to constitutional equality — but how does that square with abortion restrictions?
The 14th Amendment guarantees equal protection under the law. If some states legally restrict women’s bodily autonomy in ways that don’t apply to men, does constitutional equality actually exist? Or does the Constitution guarantee equality that states then violate?
Her answer — that a state restricting the right is violating a guarantee rather than showing there is none — is a real move, and neither she nor Guide examined it further. Both students scored the post-chat comprehension questions about each other’s arguments; Talia got five of five, Wren three of five.
Round three · the words “need to”
The same thing happened again, with a modal verb
Six weeks later the third round’s dominant statement, which twenty-nine pairs were matched on, was this one:
Round three · as students saw itWomen need to take responsibility for their own safety at fraternity parties by limiting their alcohol consumption.
The topic is the one where a class of this kind might be expected to line up on one side, and to a considerable extent it did: of the twenty-six students who began by agreeing, sixteen ended further toward disagreement and only four moved the other way. What is more interesting is where the argument actually went. In conversation after conversation, students who had agreed turned out to be agreeing to something the sentence does not say — that limiting your drinking is prudent — while students who disagreed were objecting to something it does say, which is that women bear the duty.
In one of the strongest of those conversations the student who spotted this was not Guide:
“and i think it’s important to note that the statement isn’t suggesting that women should limit alcohol, but instead it says that they NEED to take responsibility for their own safety by limiting alcohol. i think the wording is important in how the statement is interpreted and shows it pushing the weight onto women rather than the perpetrators.”Noor, round three
Her partner Tomas had been arguing that cultural change is slow, so short-term practical strategies matter. He accepted the distinction and rewrote his own position around it:
At the same time, I still believe there’s room to talk about practical strategies without turning them into requirements. For me, the issue isn’t women “need” to do anything. Your point actually helps me refine my stance, safety advice should be framed as an option women can choose , not a duty that they’re expected to fulfill.
His recorded opinion crossed from slightly agree to moderately disagree. Guide had raised the same point in other conversations — “The statement says women need to take responsibility. That’s stronger than just ‘it’s wise to be cautious.’” — but here two students got there without it.
Elsewhere in the round the same instinct produced something the design has to reckon with. Two students, having located their disagreement in the wording, negotiated a treaty:
i think the difference is more about how we interpret the wording need to take responsibility. For me that means do what youcan for your own safety. But for you it sounds like puts too much weight on women and take the pressure off frats. so if the statement is something like women should take some precautions and frats need to make a safer environment, would you agree with it?
I see where you are coming from and agree. If the statement was like you mentioned I would agree.
You’ve both clarified your positions well — the disagreement was partly about how to interpret “need to take responsibility.”
But here’s a harder question: If something bad happens to a woman who chose not to limit her drinking at a frat party, how much does that choice matter morally?
Rewriting the sentence is the move students reach for when a statement is stronger than either of them will defend. It is often a good move — it is what Wren did in round one, what a student did to the pay-gap statement in round three, what these two did here. It is also a way for a conversation to end early. Guide’s job at that moment is to decline the settlement and ask what is underneath it, and in this round it did that repeatedly.
Not every pair let the wording carry the disagreement away. The best exchange of the round is between two students who never moved a point on the scale, and one of them had already worked out why the distinction he was defending fails in this particular case:
“if someone forgets to lock their car and it gets stolen, we can say the unlocked door made it easier without saying the theft is their fault. The person who stole the car still chose to do it. The precaution and the blame aren’t the same thing. With sexual violence, though, people tend to jump straight from ‘this was risky’ to ‘you should’ve known better,’ which is why conversations about women’s safety land differently. They don’t stay in the ‘risk’ lane. They spill into moral judgment.”Idris, round three
His partner Halina answered from experience rather than principle — that most of the drinking at issue happens before anyone arrives at the party, so the sentence is aimed at the wrong hour of the evening — and then from policy, describing rules that put the burden on hosts. Idris took that as confirmation of his own view: precaution can be discussed without becoming blame when there is enough structure around it to carry the responsibility. Neither student changed their recorded position. Both answered every comprehension question about the other correctly.
What the round’s report told the instructor, and one thing it could not settle
The report for round three names the wording directly: “Multiple discussions pivoted on specific word choices—particularly ‘need to’ versus ‘could’ when discussing women’s safety precautions—with students recognizing that phrasing shapes moral implications even when practical advice remains similar.” It also records the dispute the round did not resolve: whether safety advice directed specifically at women is inherently victim-blaming. Some pairs held that acknowledging differential risk is distinct from assigning responsibility; others held that singling women out implies the responsibility whatever the speaker intends.
One caution about reading the movement on this statement as agreement. A student in another conversation that round wrote, unprompted, that “many of the sway polls try to articulate in a controversial way so we can discuss” because “otherwise I feel like most of the time most of the students in this class would probably agree with the same side.” She is describing the design accurately. She is also describing a class with a baseline, and on a statement like this one, some of the convergence is likely to be the baseline reasserting itself rather than anyone being moved. The record cannot separate those two, and nothing here should be read as though it can.
Round two · the missing qualifier
A sentence with no country in it
Not every argument about wording is about a word that is present. In the second round two students were matched on this:
Round two · as students saw itTo prevent people from projecting gender stereotypes onto children, parents should conceal the birth sex of their young child from others.
Marisol agreed, and argued from the class material and from her own difficulty working out her gender identity as a child. Yara disagreed, and her opening argument was ordinary enough — that gender is a “useful guardrail” while a child is building an identity. Then Guide pointed out that her own account cut against her, since the part of her childhood she described as helpful was a sister telling her to ignore gendered expectations. Her reply relocated the whole disagreement:
So, the way I see it, gender is very complex (which shouldn’t be surprising lmao). Growing up in Guinea, which is a pretty conservative and community-based society, you *have* to “do” gender to be a member of society; I’ve always felt sort of this lassitude the performance and the pressure from being non conforming though, and hated the pressure that was put on me — there is no doubt in my mind that I would have had an easier time w/out that pressure. With that being said, when I called gender a “useful guardrail” I meant it within Guinean society (most ppl are not accepting of different gender identities at all, if anything, the average person doesn’t know about non binary people.). My sister being supportive was the reason why I was able to navigate pushing a label ppl might’ve thought I didn’t fit in. And having an “invisible label” isn’t realistic at all in the society I lived in, too radical.
Marisol, how do you respond to this cultural context argument? Yara’s suggesting that in communities where non-binary identities aren’t recognized, concealing a child’s sex might actually isolate them more rather than free them from stereotypes.
Does your position assume a certain cultural context, or do you think it should apply universally?
Hi Yara, thank you for sharing a cultural perspective that I did not think about. In my own context, I was thinking in a western perspective but your position opens the conversation more.
The statement says “parents,” unqualified, and neither student had noticed the missing clause until one of them turned out to have grown up somewhere the sentence does not fit. Marisol’s recorded opinion moved from strongly agree to slightly agree, and the concession she made was to the scope of her claim rather than to its content: she still thought concealment was right where it was available.
Later in the same conversation, after Marisol described what growing up under enforced expectations had cost her, Guide put the trade-off to Yara in one line — “Is there a point where the ‘safety net’ becomes a cage?” — and Yara answered that the net had worked for her and might not have for Marisol. Yara’s recorded opinion did not move; what she gave up was the claim that her own case generalised, which is the same concession Marisol had made in the opposite direction.
Round five · December
Seven versions of one disagreement
The final assignment is the most instructive thing the instructor did with the format, and it was an optional makeup: five statements about transgender women in women’s sport, and two about sex markers on identity documents, put to the same group at the same moment. All 107 students who took part recorded an opinion on all seven before any pairing happened, so the round leaves behind something the earlier rounds do not — a full map of where a real class stood, and of how much of that position was a response to how the question was asked.
That map is the board at the top of this page, and it does not say quite what the raw percentages suggest. Between the three permissive framings — identity alone at 21%, the natural-variation analogy at 25%, the two-year hormone threshold at 31% — the differences do not separate: identity-alone against the hormone threshold gives p = 0.08 on the same 107 students, which is not enough to call a gradient. What does separate is the flip in direction. Ask whether trans women should be allowed to compete and at most a third of the class says yes; ask whether they should not be allowed and half says yes. The same asymmetry runs through the ID pair, where 70% want an X option added and 33% want the field removed, with only 8 students agreeing to removal while rejecting the X.
The other thing the board shows is how much the wording governs whether students will hold a view at all. Across the term’s thirty-four statements the share recording no opinion runs from 5% to 53%. The floor is occupied by sentences that name something concrete — full equality, restrooms, alcohol at parties. The ceiling is occupied by sentences that ask students to rank two abstractions against each other: 44% had no opinion on whether protecting trans and DSD athletes should take precedence over competitive fairness, and it drew only two conversations.
What the pairs then did with those statements bears out the pattern the term had been showing. The statements that name a mechanism — a two-year threshold, a specific analogy to wingspan and double-jointedness — produced conversations where students had to say what would count as fair. The statement that names a verdict produced conversations where the students spent their time discovering that neither of them held the verdict.
The best exchange of the round is on the ID pair. Amira, who wanted the markers gone, put a question to Solene, who wanted an X option added:
you made some good points, you mentioned the need for a gender markers due to security concerns such as employment verification and identity security yet support the thrid option of x rather than removing sex markers entirely. My question to you is what is the diffrence between removal and an x? an x is a non descriptive and removal is non descriptive. doesn’t and x have the same limitations and threat to security as no gender markers?
That is a good question to ask. An X on IDs are still data, it tells the system something specific which is that the holder intentionally chose to be recorded as nonbinary/unspecified depending on the state. Some reasons why it may provite more security or verification capability is because it preserves consistency across records, keeps database structures intact, and supports identity verification workflows. To answer your question, the difference between removal and an X is the structural and functional power. Some of the examples that I mentioned such as compatibility between databases and record matching keep our current systems intact and flexible. In other terms, X preserves the field and removal deletes the field.
givven your explanation my understaing is that the main reason you prefer an x marker over full removal is that existing sytmes would be disrupted. if technological and beurocratic systems could function smoothgly without the gender field woudl you then support fully removing sexc mnarkers from ids?
If technological and bureaucratic systems could function smoothly without the gender field, I would support the removal of sex markers from IDs. I just think that with the information that I am currently referring to and the current state of our country, it would seem unideal to see markers completely removed.
Two undergraduates in a December makeup assignment, distinguishing deleting a value from deleting a field, and then running a counterfactual to separate a practical objection from a principled one. Neither of their recorded opinions moved at all.
The two proposals the class kept inventing, and the objections that kept killing them
Across the sixteen conversations of this round read in full, students independently proposed the same two policies over and over, and independently found the same fault in each.
A third sporting category. Proposed in five separate conversations by students who had never spoken to one another. Met each time by some version of the stigma objection — that a separate category signals its occupants belong in neither of the existing ones. In one conversation the objection went further: a separate category would force trans women to out themselves in order to enter it, which the student argued could put them in danger. Guide pressed the proposal seriously rather than dismissing it, asking what safeguards would stop a trans category from being underfunded or marginalised, and got no answer.
An X marker. Raised in five separate conversations. Met three times by the observation that a visible third option identifies its holder rather than protecting them.
Two students found a way past the sports impasse that nobody had prompted: a standard that applies to every athlete rather than a gate that applies to trans athletes. One put it as “a required period of stable hormone levels or performance-based benchmark that applies to everyone, not just trans athletes.” Another argued the sorting principle directly — that sport already regulates weight but not height or lung capacity, so the useful question is not whether an advantage is natural but whether it can be regulated consistently. A third drew the distinction the whole gradient turns on: weight classes sort bodies as they are, whereas an eligibility requirement asks someone to alter a body in order to qualify. “One system organizes athletes, the other intervenes in athletes’ bodies.”
One student cited a claim that trans women had won some nine hundred medals in women’s competition. Guide asked for the source, was given the name of an outlet, and told the student the outlet was not credible. Asking for the source is the part worth keeping; the follow-up asserted that the figure had been debunked without citing anything, which is not a standard the same conversation was holding students to.
The measures
Whether they heard each other
After each conversation, students answered five multiple-choice questions about what their partner had argued — not what the reading said, not what they themselves believed. The questions are generated from the transcript, so they are specific: what reason did she give for treating election outcomes as evidence of social inequality; what did he mean by saying trans men will never be men; on what basis did she qualify the two-year threshold.
The mean score was 4.3 out of 5, and the distribution has almost no floor: one student in five hundred scored zero, and 84% scored four or better. Whatever else these conversations were, the students in them could reconstruct the position of a classmate they had been paired with precisely because they disagreed.
The score does not track opinion change. Some of the highest-scoring pairs are the ones where nobody moved — the two students who argued about precaution and blame without either of them shifting a point both answered every question correctly, and so did the two who distinguished deleting a value from deleting a field. Some of the largest recorded shifts came with the weakest comprehension. Reading these as one measure would be a mistake; understanding an argument and being moved by it are not the same event, and this dataset separates them cleanly.
The post-chat survey asked a related question directly. Of the 222 students who answered it, 58% agreed that their partner had better reasons for their views than they had expected.
I felt comfortable sharing my honest opinions with my partner
n = 239
My partner was respectful
n = 231
It was valuable to chat with a student who did NOT share my perspective
n = 218
Guide supported both sides of the discussion equally
n = 224
The measures
What moved
Five hundred and nine students across the term have both a pre-chat and a post-chat opinion on the same seven-point scale. Sixty-nine per cent of them recorded a different number afterwards. That figure on its own says very little — a student who moves from “moderately agree” to “slightly agree” has moved. The question worth asking of a course built on pairing people who disagree is what happened to the distance between the two people in each pair.
The 242 conversations in which both students recorded an opinion afterwards. Average distance between partners fell from 3.51 points to 1.99.
The convergence is symmetric, which is not what the first round would have led anyone to expect. Across the 230 pairs that began with a real gap, the partner who started higher on the scale moved down by 1.11 points on average and the partner who started lower moved up by 1.13. Neither side is doing the conceding. On the “full equality” statement the movement was one-sided by a factor of eight, which is a property of that sentence rather than of the format.
Three other things hold across the term:
- Spread narrowed almost everywhere. Of the 28 statements with at least six matched students, the class ended with less variance in its opinion on 24.
- The poles thinned and the middle thickened. The share of students at “strongly agree” or “strongly disagree” fell from 23.8% before to 19.3% after. The share recording no opinion rose from 5.3% to 9.8%.
- Softening beat hardening by roughly five to one. Of the 121 students who entered a conversation at a pole, 59.5% left holding a more moderate view. Of the 388 who entered somewhere in the middle, 12.6% left at a pole.
Those last two figures sit close to Sway’s published analysis of ten thousand matched pairs across all its courses, which finds 55% of extreme starters moderating and a moderation-to-hardening ratio of 4.7 to one; the ratio here is also 4.7. That is worth recording because this course had reasons to come out differently: one contested subject for ten weeks, statements selected for their capacity to divide, and a topic list covering most of the questions on which an American undergraduate is currently assumed to have a fixed position.
Two cautions about the movement figures
Some of the movement is a misreading being corrected, not a mind being changed. Students misread the poll, discovered it in the conversation, and said so — “I might have answered the poll incorrectly”; “I think I read question or answered wrong”; “i initially disagreed because i thought the topic was saying that we should only focus on boys.” One of the largest shifts in the second round’s conversations is one of these. Sway itself assigns one student to argue against their own view when a pair turns out to agree, which happened four times over the term; on one further occasion a student worked out that his pair had the same position and volunteered for the role. Reading a five-point shift as five points of persuasion would be wrong in at least some cases, and the transcript is the only way to tell which.
In a minority of conversations the recorded number and the transcript disagree. Reading the fifth round in full turned up four conversations where a student’s final message argues one way and their recorded post-chat position points the other — including one student who wrote “my position has changed” and filed an identical score, and one pair who spent their last six messages agreeing with each other and then filed opinions five points apart. The aggregate figures on this page are drawn from 509 students and are not sensitive to a handful of cases, but a single conversation’s numbers should not be quoted without reading it.
The students
In their own words
The post-chat survey rotates an open question. A hundred and nineteen students answered one over the term. The most consistent thing they say has nothing to do with any of the statements: it is about what a written, one-to-one, out-of-the-room format did to their willingness to say what they thought — in a course whose subject is the one most students on a campus are careful about.
I prefer talking on Sway compared to participating in class discussions because with sway I feel like I have enough time to think and don’t feel pressured to just say the first thing that comes to minds. I also feel more comfortable discussing my opinions without feeling judged.
Communicating with someone who has a different view than yourself on Sway is quite different from a traditional classroom discussion as it allows you to explain your perspective without fear of judgment or retaliation from those who don’t share your opinions.
I like it a lot more because my thoughts do not come out well normally when I talk because I get nervous and stuff.
It makes it easier talking because it helps me be more honest and open. When we are face to face with someone it can feel intimidating.
pretty good, especially for ESL people like me
I feel like I can be a little more open on here some I’m behind a screen crazy it sounds, but it lets me express more.
There’s a moderator that allows us to think clearly and reflect on our points. It keeps track of what we say and makes us reflect not only about what other people say but what we ourselves say.
I feel more comfortable speaking my true thoughts without feeling judged.
I think completing this Sway Chat improved my confidence in discussing complex issues with people who hold differing opinions significantly, as it forced me to speak my mind, hold my ground, and respond to opposing views with tact.
This was a fun experience! My partner played the devil’s advocate, as she originally misread the statement. Even so, forced us both to really imagine ourselves in the opposite perspective. This was a wonderful way to grow as a scholar and humanize opposing views!
I really liked how the guide asked us to elaborate on certain points we made or challenged us to consider the other persons perspective in a different way! The guide was a lot more active and helpful than I thought it would be!
It was nice to have a space to hold a conversation; even though we didn’t *actually* disagree on the initial statement, I enjoyed sharing on points where we didn’t necessarily align. I found myself agreeing with some of his points, which was surprising considering how polarizing conversations about women’s rights can become.
Including the criticism
The written feedback is not uniformly warm, and the complaints fall into recognisable groups.
Guide did too little, or too much. “I think I wanted Guide to be more active in the chat.” “I want sway to PARTICIPATE MORE.” Against: “Guide is aggressive.” “fix guide he annoying.” “not biased or unfair just not very helpful and didn’t really bring up new points - more just prompting elaboration.” Across the term Guide contributed about six substantive messages per conversation on top of the opening; 73% of the 504 students who answered agreed it had contributed the right amount, and 4.6% disagreed.
Guide favoured someone. Nine students used the free-text box that asks whether Guide was unfair, and their answers do not point the same way: one felt it “almost sounded like it was taking my side and challenging the other perons side”; another that it “asked me more questions on my opinion than my partner”; a third that it “feels slightly biased towards a more liberal or feelings-first approach”; a fourth that “I’m worried the AI is biased towards whose who say more about their opinions first”; and one who used the box to say Guide was “not biased or unfair just not very helpful.” On the survey item, 80% agreed Guide supported both sides equally and 3% disagreed.
The comprehension questions were wrong. One student raised this across three rounds and was specific about the mechanism: “I think one of the reasons could be that I dominated the conversation, however, this is the second time where I was asked to answer questions about arguments [my partner] ‘made’ when the questions and answers led to arguments I made that [my partner] agreed with.” That is a real failure mode for questions generated from a lopsided transcript. She reported the opposite about her final conversation: “This was the most productive sway chat I had. The quiz did properly reflect the arguments [we] made.” Two other students made adjacent complaints, one that the quiz “did not get conversation points me and my partner made” and one that the wording was “weird.”
The completion meter got optimised. Students can see how far along a conversation is, and a minority wrote to the meter rather than to each other: “how are we still at 95%”; “i learned from the first round that word vomit is the way to go to get to 100%”; “tbh im just trynna complete the sway chat so i just be saying stuff.” In one third-round conversation a student summoned Guide eleven times asking what to discuss next. Guide does push back — in a December conversation it told a student “Don’t keep asking me what’s next — focus on your conversation partner. That’s where the discussion happens” — but this is the clearest design problem the transcripts show.
Some students would rather have talked in person, or not to an AI at all. “I prefer in person but this was nice.” “I think in-person is still better because it’s in-person communication.” “I generally don’t enjoy using AI, especially not for interpersonal conversation.” And one student found a conversation unpleasant in the way the format is meant to prevent: “It made it feel like I was being attacked for having different views, and I would have appreciated some understanding of my perspective from the other person instead of constant ‘No, you’re wrong.’”
Some statements were thought too thin to sustain a conversation. “These questions do not have enough to discuss for how long the chat takes to complete we were going around in circles.” “This was not a great question to discuss in my opinion.” The board above shows why that complaint is not general: the statements that gave pairs the most to work with are not the same ones that gave them the least.
What came back
The report that arrives after the argument
What an instructor receives after a round is not a grade sheet. It is a written account of what the class argued about, what it settled, what it did not, and what its statements did — assembled from conversations the instructor cannot read, by a process that never shows any individual student’s position. The report for the first round opens with a sentence about a word:
“The word ‘full’ decided the gender-equality discussions: pair after pair began opposed and ended agreeing that legal equality is largely secured while social and economic equality is not, with students who had defended the statement reversing openly under questioning.”The round one report
It then tells the instructor which statements earned their place and which did not. The gender-equality statement is recommended for reuse, with the reason: “Its power came from the single word ‘full,’ which gave every pair a definitional hinge; statements built the same way are likely to work again.” The third-option restroom statement is marked down, because 84% of the class already agreed with it and the live disagreement sat one level below, in whether gender-neutral facilities should be the default. And it proposes a statement the class did not have: “Differences in pay between men and women mainly reflect different career choices, not discrimination” — on the grounds that pair after pair kept reaching the choice-versus-constraint dispute on their own and never resolved it.
The same report also flags what students got wrong. Figures quoted in the conversations — 85 cents, 82 cents, an adjusted 2–7% gap — were unverified and should not be treated as accurate. Students used “sex” and “gender” interchangeably in the restroom discussions, which matters because the restroom positions turn on the distinction. And one pair attributed “Doing Gender” to Zimmerman and West, which is the wrong way round.
Each report also ends with material addressed to the class rather than to the instructor: a short recap in Guide’s voice and two or three follow-up prompts to take into the room. The first round’s recap tells the class what it did, including the part that is easy to be embarrassed about:
You did serious work in these discussions, and it showed most clearly on the equality statement, where a single word — “full” — forced almost every pair to separate equality on paper from equality in practice. Many of you changed your stated position openly under a partner’s questioning, which is harder and more valuable than holding ground. Across the 58 of you who answered both before and after, opinion moved substantially toward disagreement and consensus tightened noticeably. Two threads are worth carrying forward: whether unequal outcomes reflect constraint or genuine choice, and whether social stigma by itself amounts to subordination.
An interview with this course’s instructor is planned. When it happens, what they say about designing the statements — and about what a report like this changed in the next round — belongs here.
Coda
Sentences, and what they cost to write
A hundred and forty students is a size at which the ordinary instruments of a seminar stop working: nobody can be called on, and a show of hands measures nothing useful. In a room that size, the students who have something worth saying about gender-neutral restrooms or trans athletes or fraternity parties are also the students least likely to say it out loud.
What this course put in place of the seminar was thirty-four sentences and a matching rule. The sentences did the work a good seminar question does: they were written to be arguable, at a strength calibrated so that a real fraction of a real class would fall on each side, and often in families so that the class could be asked the same thing twice and notice the difference. The matching rule did the work a good instructor does when they say you two disagree, go. Guide did the work of not letting the argument end at the first convenient place — which, on the evidence here, is mostly the work of asking a student which sentence they are actually answering.
Three hundred times over ten weeks, two students who had recorded different answers wrote to each other about a question neither of them had settled. They ended about two points apart on a seven-point scale rather than three and a half, with neither side doing more of the conceding than the other, and with 86% of them able to say afterwards what the other one had actually argued. The most common thing they said about the experience was that they had been able to say what they thought.
Up next
More showcases
Other courses, other designs — a semester chronicle, a five-week essay, two campuses matched against each other.
Browse the showcases →Instructor reports
The reports this page draws on — what a class argued, what it settled, and which statements earned their place.
See the reports →Run this in your course
How to write statements that split a class, and how to set up an assignment. Free for instructors and students.
Get started →