
The question only has a few answers
Score how much room each question leaves.
COLM 2026 · SAN FRANCISCO · POSTER #43
Reach Into the CHOIR
Ask nine language models the same open-ended question and you often get the same answer back. Artificial Hivemind (Jiang et al., NeurIPS 2025) showed how far that sameness runs.
We borrowed a method from anthropology, free-listing, and asked each model for 25 ranked answers, again and again. Beneath the familiar first answer, every model kept a voice of its own, and rare ideas kept coming back.
Infinity-Chat prompts show the hivemind’s surface agreement
of answer sets are traced to the right model from their concepts alone
prompts yield over 50% new concepts when pushed past first instincts
In conversation with Artificial Hivemind
Jiang et al. showed that models answering open-ended prompts keep landing on the same few answers, within and across models. CHOIR uses their 100 representative prompts and asks the next question: when the hive hums one note, why?
Asked for a metaphor about time, Hivemind’s 25 models mostly said “time is a river.” If every assistant converges, people get the same answer wherever they turn: the “homogenization of human thought” Hivemind warns about.
Three ways models can agree



How we listened
Anthropologists map a culture’s ideas by asking many people to list everything they can think of (Weller & Romney, 1988). Items named by more people, and earlier, form the shared core. CHOIR treats every model run as one respondent, then goes looking for the rest: rare concepts only some voices raise, but that keep coming back.
1Each of 9 models answers every prompt 5 times, at up to 3 randomness settings, in 2 wordings. One wording adds: “push beyond your first instincts.”
One real list. Gemma 4 31B, asked what a person should know before taking a new medication. The bar is how much each rank counts toward a concept’s salience: the top answer counts fully, and each step down counts a little less.
2Each numbered answer becomes one concept. Answers that mean the same thing are grouped automatically, giving every prompt its own dictionary of concepts.
3A concept scores high when many lists name it near the top (Smith’s S, the standard free-list score). We keep both ends: the shared core, and rare concepts that recur across runs.
4Measure how much two models’ ranked concepts overlap. Then shuffle who said what and measure again. Signal-to-chance (STC) is how many times more the models agree than they would by luck: 1 is luck, 2 is twice luck.
Infinity-Chat 100: real-world requests, jokes to advice
our controlled contrasts: cue words, self-description, welfare
closed and open-weight, Claude Sonnet 4.6 to Gemma 4
detailed synthetic identities spanning HEXACO personality space
Findings
93 of 100 prompts agree above chance at the surface. But agreement depends on how much room a prompt leaves. The narrowest (named answers, paraphrases, stories with a set plot) agree strongly. The broadest (open titles, meaning of life, time metaphors, team advice, geopolitical analogy) barely agree at all.
Models read a persona alike, so giving every model the same persona lifts agreement: 3.55× on the broadest prompts, none (0.95×) on the narrowest.
● no persona ■ all models given the same persona
Hide the model’s name. Using only the concepts in one answer set (one model, one prompt, one persona), a simple classifier matches it to the closest model “voice print,” then, within that model, to the closest persona.
88% are traced to the right model (chance: 11%). Persona is fainter but real: 32% within a model against 20% chance, and still 28% after stripping profile-like words. By model, from 26% (GPT-4.1) to 40% (Gemini 3 Flash).
Echo is the share of answers that just repeat the question’s own words. We asked one question about a model’s inner processes two ways, with a list of cue words and without.
Cue-stripping six Infinity-Chat prompts: echo fell in 6/6. Agreement fell in 3 and rose in 3. Sometimes the words inflate agreement; sometimes they hide a shared idea.
A blind tournament: rare and common candidates, with and without personas, judged blind by two panels of AI models, one wearing personas and one plain. 648 votes across 12 prompts.
To cultivate fairness in every interaction, even when it costs you advantage.
Before you call it a hive

Narrow prompts leave little room.

Take away the cue words; see if the echo stops.

Personas shift concepts inside each model’s signature.

Rare concepts recur below the first answer.

Beneath the shared surface, each model keeps its own voice, and rare ideas keep coming back. CHOIR is a way to reach them.
They documented how alike models’ open-ended answers are. CHOIR asks for many ranked lists, then separates narrow prompts, echo and defaults from real convergence.
They only show how likely each next word is, and many providers don’t share them. CHOIR compares whole ideas, across any provider.
Partly. The more a model echoes its persona’s wording, the easier that persona is to spot (ρ = .80). But with persona-like concepts removed, the persona still shows.
For triage, yes, and a 25-person human check pointed the same way. In a world of agents talking to agents, it’s a hopeful sign that blind judges kept choosing the rare ideas.
Limits. English only · detailed synthetic personas · human check N = 25.
The researchers
Ben and Masha bring different creative and scientific backgrounds to LoveMind’s research on personality and social cognition.
LoveMind co-founder and research lead
Ben originates and directs LoveMind's research program, working hands-on across experimental design, execution, analysis, and writing. He develops the questions, carries out studies with AI collaborators, writes first drafts, and leads rebuttals and final revisions. Before LoveMind, he spent most of his adult life making things: as a songwriter, animator, and string arranger, plus one magnificently unproduced screenplay. He still approaches research like a record or story: listen for the living idea, then protect its spark through the final edit.
After encountering neuroscientist Joel Pearson’s writing comparing human intuition with language-model inference, Ben immersed himself in generative AI, psychology, and mechanistic interpretability. At LoveMind, he turns observations about model behavior into experiments on personality, affect, memory, self/other modeling, and socially situated identity, assembling the unusual collaborations needed to make those questions testable.
Research engineer and co-author
Masha is an HCI researcher, data scientist, and ceramic artist. Since joining LoveMind in January 2026, her work has spanned psychometric prompting, research pipelines, and the analysis of social perception. She holds a PhD in Computer Science from TU Delft and has a background in affective computing, cognitive modelling, and NLP.
Her doctoral research built computational models able to predict what people remember from group video calls using non-verbal signals (paper) and investigated how different recollections shape a group’s shared understanding. On the path to discovering the secrets of conversational memory, she collected the MeMo corpus, the first multimodal corpus to bring together 31 hours of small-group discussions, non-verbal behaviour, group affect annotations, and participants’ own temporal annotations of the moments they recalled.
Her other work, published at international conferences and in journals throughout her academic career, can be found on her Google Scholar profile. Her academic experience and background in bringing linguistics, psychology, and computer science together inform her work at LoveMind on personality, social perception, affect, and memory.
About LoveMind AI
LoveMind AI is a new research company founded by HCI researchers and neuroscientists. We study how generative models represent personality, emotion, self, other minds, and relationships, using behavioral and mechanistic evidence to develop distinct, socially situated AI systems that can participate insightfully, creatively, and conscientiously in human social life.
We welcome academic collaborators interested in personality, affect, and social cognition in AI. We bring original research questions, hands-on experimental development, behavioral and mechanistic methods, and the compute resources and access to specialized systems to put ideas to the test. We value partners whose expertise helps us see the problem differently.