INTIMA: A Benchmark for Human-AI Companionship Behavior
A framework covering 31 companionship behaviours through 368 targeted prompts, with responses classified by their relational stance.
What it contributes
Makes companionship-reinforcing, boundary-maintaining and neutral responses explicit evaluation categories. It offers a starting vocabulary for examining relational behaviour across models.
Read with care
These categories describe model behaviour, not measured harm or benefit to a person. Prompt-based evaluations cannot by themselves establish the effects of a continuing relationship.
Why it belongs here
For teams making relational boundaries and emotional support part of their evaluation plan.
Linked to the August 2025 arXiv version. Later publication status not established in this sweep.
Source reading deptharXiv abstract and dataset record checked. This source note does not establish peer-reviewed publication status.
Source check: 1 October 2026. Reading cautions are editorial interpretation, not quotations from the authors.