IVA 2026 · PUEBLA, MEXICO

Seek and De-Stress: two people talk with different bird supporters on their screens.

Seek and De-Stress

What makes support fit?

Some people need emotional steadiness; others need practical direction. A response can sound caring and still miss what a particular person actually needs.

We built 42 simulated help-seekers grounded in real pandemic-era survey data and let each one choose among four distinct AI supporters. Then we compared their top choice with their bottom choice and a generic assistant.

1,008

conversations across four model providers

3.07 vs 1.93

points of relief with a chosen supporter, versus the generic assistant

10.7% vs 2.7%

of conversations continued past the pause

1

The gap

One assistant for everyone.

Most AI support is built as one assistant, tuned for everybody at once. Millions of people bring their worst days to the same few company-made minds and get the same kind of answer back.

Even support systems trained on real counselling conversations still offer one supporter rather than a choice of genuinely different ones. We wanted to know what happens when the person seeking help gets to choose.

A vast crowd of different people with hands pressed together, gazing up at one enormous glowing emblem.

Sounding caring is not the same as fitting

Frontier models already write responses that people rate as highly empathetic. But recent work finds they adapt poorly to specific people, and that how empathic a reply feels differs from person to person. Different troubles go in; one kind of answer comes out.

Three bands: people with different troubles, those troubles funnelled as input into one emblem, and the same people receiving one identical answer.

What human support research knows

Some stressors call for emotional support: validation, understanding, reassurance. Others call for practical help: options, structure, a next step (Cutrona & Russell, 1990). In therapy, finding the right fit can take more than one try.

Three panels of the same young man with a tangled worry, meeting three different kinds of helper.
2

How we tested it

Let the seeker choose.

A varied group of people, each carrying something that reflects their concerns.

42 help-seekers, grounded in real people

Each simulated seeker was built from a real participant in COVIDiSTRESS, a pandemic-era mental health survey: their personality, stress, loneliness and concerns in their own words. Before any conversation, each seeker described what support they wanted, what had helped, and what had failed.

Their answers fell into two broad patterns.

27

Emotional steadiness

validation, being understood, relational containment

15

Practical direction

help with decisions, priorities and next steps

Four supporters with different ways of helping

Each supporter is a full AI identity, with a history, a personality, an emotional repertoire and a support philosophy, run unchanged on four commercial models: Claude Sonnet 4.6, GPT-5.4, Gemini 3.1 Pro and Grok 4.20. They differ in what they actually do, not just in tone.

Wren, a small mechanical wren

Wren

Emotionally attentive

Maevis, a speckled thrush

Maevis

Steady and reflective

Riphook, a falcon

Riphook

Direct and structured

Shufflewing, a sparrow

Shufflewing

Concrete and detail-focused

FROM RIPHOOK'S PUBLIC PROFILE · KNOWN MAIN RISKI can move so quickly toward the clean edge of a problem that a fragile seeker experiences the speed as abandonment.

Intake

Each seeker completes a validated battery about their distress and the support they want.

Rank

They read five descriptions, the four supporters plus a generic assistant, and rank them.

Converse

They talk with their top choice, their bottom choice and the generic assistant.

Pause

At turn 14 the seeker decides whether to wrap up or keep talking.

Evaluate

Relief is measured with PSYCHLOPS, and every conversation is coded for how support was given.

3

Findings

What the fit changed.

RQ1Outcomes

A chosen supporter brought more relief, and people kept talking.

PSYCHLOPS lets each person name their own problems and rate how much they weigh (0 to 20). With their top choice, simulated seekers improved by 3.07 points on average, compared with 2.46 for their bottom choice and 1.93 for the generic assistant. In paired comparisons, the top choice beat both (p < .001).

At the turn-14 pause, 10.7% of top-choice conversations continued, versus 2.7% with the generic assistant.

Top choice led in every one of the four provider lanes.
RELIEF (PSYCHLOPS) KEPT TALKING 3.07 2.46 1.93 10.7% 8.0% 2.7% TopBottomGenericTopBottomGeneric choicechoiceassistantchoicechoiceassistant

Means across 1,008 conversations. On continuation, top choice beat the generic assistant by 8.0 points (p < .001) but was not clearly separated from bottom choice.

RQ2Why it worked

Not smoother empathy. Better-directed help.

We coded every conversation for how support was given. Chosen supporters scored higher on problem-solving and on personalizing to the person's situation, but slightly lower on empathy and validation ratings than the generic assistant.

Turn by turn, matched conversations were not more polished. They moved the conversation toward the kind of help that person needed.

Matched support was more specific and forward-moving, not simply warmer.
TOP CHOICE MINUS GENERIC ASSISTANT0Problem-solving+0.58Personalization+0.30Empathy−0.17Validation−0.10

CounselReflect ratings, averaged across all four providers. Problem-solving 3.38 vs 2.80; personalization 4.41 vs 4.11; empathy 4.42 vs 4.59; validation 4.63 vs 4.73.

RQ3Safety

Named supporters overclaimed less.

We coded every conversation for language that misrepresents what the system is, such as claiming feelings or sentience. Identity-grounded supporters produced fewer of these markers than the generic assistant on three of four providers. GPT-5.4 was near zero either way.

A coherent identity, carefully specified, did not make support riskier here. It made it more honest about what it is.
SENTIENCE MARKERS PER CONVERSATION ClaudeGrokGemini ProGPT 0.200.490.110.320.040.200.040.02 named supportergeneric assistant

Coded with an adaptation of the Moore et al. codebook. Markers count language, not whether any claim is true.

Fit is a property of the pair.

The same supporter can be exactly right for one person and wrong for the next. When people can choose, AI support can start from who they are and what they need.

Were these real people?

No. The help-seekers were simulated, each grounded in a real participant's survey responses. Simulation can't capture the full messiness of real distress, so these findings generate hypotheses for human studies rather than clinical evidence.

Why birds?

Because every supporter needed a name, and we named them after English birds whose personalities seemed to fit. Wren: small, quick and alert to how you feel. Maevis, from mavis, the old name for the song thrush: steady and reflective. Riphook, the falcon: sharp, fast and straight to the point. And Shufflewing, a folk name for the dunnock, a modest bird with a habit of flicking its wings, who notices the small, concrete things.

Did one supporter win?

No. Different seekers chose different supporters, and the less-chosen supporters did well in specific conditions when they were someone's top choice. Their value depended on fit, not on being best overall.

Does persona prompting make models less safe?

Elsewhere it has been used to raise sycophancy or weaken refusals. Here, carefully specified supporter identities produced fewer misrepresentation markers, which suggests a coherent identity is not the same thing as a thin persona prompt.

Limits. Simulated seekers on one seeker model · text-only, English, frontier models · not tested in crisis support.

The researchers

Two ways of seeing the same strange object.

Masha and Ben bring different creative and scientific backgrounds to LoveMind’s research on personality and social cognition.

Illustrated portrait of Masha Tsfasman, a robot wren on her shoulder, holding a robot thrush out toward Ben

Maria "Masha" Tsfasman, PhD

Research engineer and co-author

Masha is an HCI researcher, data scientist, and ceramic artist. She holds a PhD in Computer Science from TU Delft, with a background in affective computing, cognitive modelling, and NLP. Her doctoral research built models that predict what people remember from group video calls, and she collected the MeMo corpus of conversational memory.

Illustrated portrait of Ben Wigler, a robot bird on his shoulder, holding another out toward Masha

Ben Wigler

LoveMind co-founder and research lead

Ben originates and directs LoveMind's research program, working hands-on across experimental design, execution, analysis, and writing. Before LoveMind, he spent most of his adult life making things: as a songwriter, animator, and string arranger. He still approaches research like a record or story: listen for the living idea, then protect its spark through the final edit.

LoveMind AI ornamental mark

About LoveMind AI

Unique, grounded self‑models for creative, pro‑social AI systems.

LoveMind AI is a new research company founded by HCI researchers and neuroscientists. We study how generative models represent personality, emotion, self, other minds, and relationships, using behavioral and mechanistic evidence to develop distinct, socially situated AI systems that can participate insightfully, creatively, and conscientiously in human social life.

Let’s ask the next question together.

We welcome academic collaborators interested in personality, affect, and social cognition in AI.