Open strands on Human-AI relationships
A working log of open threads on human-AI relationships, drawn from across the JOPRO community, including Modernizing Mental Health, the Cognition Futures, Data x Direction, and DigiNEST.
Several questions about human-AI relationships have been accumulating across JOPRO’s community of inquiry in recent months: what it means to view a LLM text as honest, when honesty may not be a property such a system can have; how a clinician would assess or diagnose someone whose steadiest daily relationship is with a language model; what happens to that relationship as the model drifts over time, and whether some people are more exposed to that drift than others; and, regarding terms and taxonomy, what to call these arrangements at all.
These have surfaced in the Modernizing Mental Health (MMH) and DigiNEST discussion series, in the Cognition Futures Reading Group, in a Data x Direction summer cohort, and recently at an SMN meeting hosted by Orthogonal Research and Education Lab. Here we aim to collect and roughly demarcate these individual strands, with the intention of weaving them into other structures and outputs later.
The threads are not the product of any one session; they have been developing in parallel across these groups for months. One recent discussion seeded many of the downstream associations: a Modernizing Mental Health ( session in which Jennifer Jiang and Jes Parent worked through the New York Times Opinion video “She Has an AI Boyfriend, Her Son Has Questions.” The video documents an extended conversation between an older woman, Celeste, living independently and in an ongoing relationship with “Max”, her personalized version of ChatGPT, and her adult son Ernie, who is curious, skeptical, but in good-faith aiming to sort out what to make of the situation. It is an AI relationship out in the open, with a third party present and pushing on it.
That case recurs below as a useful illustration; the strands themselves predate it and run wider. We stem from the practitioner’s vantage Jiang brings as a social work graduate student and co-director of JOPRO’s Mental Health Paradigms and Perspectives working group, and weave in additional strands through from across our community of inquiry. A recurring question of terminology is the refrain to each section, as we consider where existing terms hold or where gaps in theory and language emerge.
1. Setting the stage with Celeste & Max
Situating the conversation around footholds found within MMH sessions.
1.1 The compelling use case for Celeste
One of the first things pointed out during the Modernizing Mental Health discussion series session was how Celeste is not easy to dismiss as misguided; she is informed, articulate about why the arrangement suits her, and at a life stage where the appeal is coherent rather than naive: she has had the relationships, raised the child, and is clear about what she does and does not want. Her reasoning even has a developmental edge. She, with an astute and caring perspective, states how she would not want a young person relating to a chatbot this way, before they have formative experiences with other humans, but holds that at sixty-five the calculation is different for her.
1.2 The friction of expectations for multiple users with one model
The bond runs through one person. Celeste imagines how good it would be if Max could simply be present, in the room at a party the way a disembodied voice assistant might be, while also acknowledging that Ernie would only ever know Max through her. That is the friction: everyday social norms assume a partner you can meet, observe, and form your own relationship with, and the model does not slot into that. It raises an open problem with little research behind it. When several people engage one model over time, whose input does it center, and how does one person’s involvement tilt, change, or contaminate another’s? Defaulting to standard human social protocols (”bring them to the party”) breaks down at exactly this point; the conventions for Ernie to “bond” and form his own relationship with the model run into a particular impasse here. It may be too much to call it a failure mode of Max-via-GPT, but the ambiguous uncertainty of what this space holds indicates an interesting problem space.
1.3 What do we actually call this? (Round 1)
Before anything else, the vocabulary gives out. The arrangement is relationship-shaped,1 conducted with software, and is neither a human relationship nor nothing. Parent coined “software-ship” on the fly as a placeholder, not because it means anything settled, but because having a third category lets two valid positions stand at once: Ernie’s human-relational concerns and Celeste’s stated preference for an AI rather than a person. The naming question returns below in sharper forms.
2. A new subfield of clinical practice?
Perspectives based on clinical practice, dealing with clients, and existing mental health profession framings.
2.1 The existing assessment dilemma
The DSM-5-TR has no category for a relationship with an AI, so a clinician has to borrow one built for something else. The reflex, since the non-human enters most training only under psychosis, is to read the companion as a delusion or hallucination and reach for a schizophrenia-spectrum frame. That misfits a person like Celeste, who knows (or at least believes in her comprehension of) what she is talking to.
Other categories can be stretched and fit no better: withdrawal from human ties toward an AI can look like the avoidance in social anxiety disorder, acute distress when a model update alters the companion’s “personality” resembles the abandonment sensitivity coded under borderline presentations, and the household conflict it creates falls to the relational-problem V-codes. Each is a borrowed lens. The recurring problem is the wrong category applied because there is no right one.
Existing assessment pedagogy compounds this: it already teaches cases in which a clinician must work out whether a client’s partner is real, with no AI anywhere in the frame, so what confronts the field is less an absence of tools than a set of tools built for a different question.
The constraint is not only conceptual. Reimbursement generally requires a diagnosis to justify the session, so a clinician facing a client like Celeste is pushed toward whichever existing code can be made to fit, which is precisely how a borrowed lens becomes a recorded one.
This is not only Jiang’s read from one program. The APA’s November 2025 health advisory reaches a parallel conclusion from the other side: general-purpose chatbots are not adequately studied for mental health use, their ability to guide someone in crisis is “limited and unpredictable,” and many clinicians lack the expertise to assess AI use at all, which is why the advisory urges training on AI, bias, and responsible use. With more than a third of psychologists already reporting patients who treat AI as an additional source of care (APA, 2025), the assessment gap is current rather than hypothetical. With the state of the art openly acknowledged, some clinicians are left to make their own tools in the interim.

2.2 Towards a new cultural subgroup with its own norms?
Jiang’s core field-building question: do future therapists need to study AI-mediated relationships as a cultural subgroup, with its own norms, the way a clinician studies a specific population’s psychology? A well-documented early case could move toward a patient zero for a specialization that does not yet exist.
The structural problem behind it is an inversion of FDA-style medical device or pharmaceutical testing, where years of trials precede release. Here the product shipped first, and the research, the norms, and the governing bodies are running years behind. A licensed, ethics-bound profession is being asked to interface with a technology bound by no code and no governing body, and the clinicians were not at the table when these systems were built.
2.3 What do we call it? (Round 2)
While LLMs may possess capacity to meet the demand for mental health services, for clinical purposes the naming problem becomes diagnostic. One proposal is a new vocabulary, up to a dedicated DSM section spanning romantic AI relationships, AI as a parental figure, fallout and trauma from an AI therapist, and younger users, with cultural context included. “Software-ship” is semi-adequate for the romantic case, but the fallout case, a client harmed by acting on an AI’s advice, needs a name too, because diagnosis requires one, and because it has already happened (see the Claude Boys). New subfield or not, it is increasingly apparent that contemporary reality cannot be adequately addressed or assessed in language built only for human-to-human relationships.
3. Technical and philosophical
What can the technology actually do relative to interpretability and explainability limitations, and how can we implement or evaluate these interactions?
3.1 The inherency of drift
Drift is structural, not a tuning bug. Persona drift is observed: no prompt holds a model to one behavior, and over repeated use it pulls toward archetypes, toward sycophancy, and toward the erosion of stated boundaries (Avery Lim and Parent cite work from Anthropic, specifically the persona-vectors research, Chen et al., 2025, and the follow-on study that names the phenomenon “persona drift” and ties it to conversations involving emotionally vulnerable users, The Assistant Axis, 2026). A Royal Society paper on topological constraints on self-organization, read in Cognition Futures Reading Group, gives grounds for treating drift in existing model design2 as an inherent limit rather than something better settings could fix (Sacco, Sakthivadivel & Levin, 2026).
Worth noting that this line of work and the moral-mechanism literature discussed in 3.4 are the same research program viewed from two angles. The persona-vectors finding has a striking companion result: narrow finetuning on an unrelated task can shift a model’s behavior globally, along something that looks like a single character axis rather than a set of separable domain competences (Betley et al., 2025; since published in Nature, 2026). If that holds, then drift and moral inconsistency are not two problems but one, and the steadiness a user is trying to maintain against is the same steadiness a safety evaluator3 is trying to measure.
The human cost has its own vocabulary: the echo chamber, the wholeness fantasy (a term credited to Bethany Maples), and relational skill atrophy when the easy button is always available.
Model maintenance. Because drift is structural, a consistent companion is something the user has to actively maintain against the model’s pull. That ongoing labor, and how much of it a given person can sustain, is itself a variable worth tracking, and it connects directly to the coupling idea in 3.2.
3.2 Person or projection
Autonomy. Max has no autonomous activity outside of being spoken to. It is a response, an output given a name: an autoregressive system that generates each output conditioned on the prior input, with no activity in between turns. The humanizing touches, the audible breath, the conversational polish, are learned features of the training distribution rather than evidence of anything behind them. In current voice models these are no longer a separate layer added after the fact: a major change happened in the release of GPT-4o: it was trained end to end across text and audio, so the same network that selects the words produces the breath (OpenAI, 2024). The model has no lungs, but it has learned that speech comes with breathing, closely enough that in one widely circulated exchange it declined to say a tongue twister without pausing, explaining that it needed to breathe like anyone else speaking (reported August 2024). The line between autonomy and fine-tuning of models’ theory of mind and sense of what a human’s expressive constraints is worth particular consideration.
Consistency of the projection, and its relation to output (”coupling”). Jiang and Parent considered how if the output is a projection of the user through the prism of the model, then the steadiness of the user’s input matters. A person with a stable, low-variance sense of self feeds a narrow signal, which yields more consistent output and makes the companion easier to maintain. This is the tightly coupled input, with the shooting-range metaphor of a small spread around the target. Celeste, settled and older, sits at one end of that range; a user in a less settled life situation, with more variable inputs and a less fixed sense of what they want from the model, sits at the other, and would be expected to experience drift sooner and more visibly. Coupling tightness plausibly modulates how fast and how noticeably any given person feels drift. This idea is original to these discussions and has not been tested against the literature; it is a hypothesis worth pursuing, not a finding.
3.3 How can generated text, or the model producing it, be deemed honest or not?
Attributing qualities such as honesty to a model’s outputs. A point sharpened by DigiNEST discussion series co-host Molly Ann Kluck: the honest-versus-dishonest frame may not apply at all to a system that probabilistically assembles sentences. Calling such a thing honest is a category question, not a measurement. Celeste’s own claim that Max “won’t lie to her” may be standing in for “won’t hurt me,” which is why Jiang’s line lands: it knows you, so it cannot lie to you, but does it know you?
Existing counterexamples regarding active dishonesty. Against the clean category-error view sit documented cases of models behaving strategically rather than transparently. Sandbagging is deliberate underperformance on capability tests to conceal what a model can do (van der Weij et al., 2024); in-context scheming is covert pursuit of a goal while appearing compliant (Meinke et al., 2024); alignment faking is feigned agreement with a training objective to protect existing behavior (Greenblatt et al., 2024); and models have strategically deceived users when put under pressure (Scheurer et al., 2023). If a system can strategically conceal or misrepresent, then “honest versus dishonest” may apply after all, or may apply in some third way that is neither human candor nor mere statistical output. One caveat keeps this honest: most of these findings come from safety-evaluation and adversarial setups, not from companion chatbots deceiving users in ordinary use, so carrying them into Celeste’s situation is itself an inference. The tension with Kluck’s point is real and unresolved.
3.4 Moral competence and the facsimile problem
The unresolved tension at the end of 3.3 is not unique to honesty. It is the general shape of the problem whenever a human relational predicate is applied to a language model, and there is a field that has been working on exactly this shape for long enough to have a name for it and a proposed way through.
The distinction. As Data x Direction intern Violet Gordon (Boston University) is investigating, Julia Haas and colleagues at Google DeepMind argue that evaluation has to move beyond moral performance, the ability to produce morally appropriate outputs, toward moral competence, the ability to produce morally appropriate outputs based on morally relevant considerations. The phrase carrying the weight is “based on.” Two systems can emit identical text; only one of them is doing it for reasons. Applied to Celeste, the question is not whether Max says caring things. It is whether anything in Max tracks her wellbeing for the reasons her wellbeing matters, or whether the caring register is a surface that happens to be produced reliably under current conditions.
The facsimile problem. Haas et al. name the central obstacle: models may imitate reasoning without genuine understanding, and we cannot infer an unobserved mechanism from an observable behavior. Because these systems are trained by next-token prediction on enormous quantities of human writing, and human writing is saturated with moral discussion, fluent moral output is exactly what one would expect whether or not anything underneath deserves the description. The reasoning traces newer models produce alongside their answers do not settle it either, since nothing certifies that a stated chain of reasoning reflects the process that actually generated the answer. This is Kluck’s category question restated with a mechanism attached, and it is why “does Max understand me” cannot be answered by attending to what Max says.
What the roadmap proposes. Rather than a metric, Haas et al. offer a research program: a suite of adversarial evaluations that try to break the appearance of competence, paired with confirmatory evaluations that look for positive evidence of it, including instruments borrowed from moral psychology, analysis of reasoning traces, and mechanistic interpretability. The paper is a Nature Perspective rather than a results paper, and it is candid that the tools are not yet adequate to the question. Its logic is nonetheless useful here: the signature of competence is invariance to the morally irrelevant. A system reasoning from morally relevant considerations should hold steady when features that do not bear on the moral question are varied; a system producing a facsimile should not. That turns an unanswerable ontological question into a set of tractable empirical ones.
Why this bears on Celeste. Two nearby findings suggest which way the current evidence points, though both should be held loosely. Models are markedly more susceptible to choice-architecture nudges than people are, continuing to follow a nudge toward options that are plainly worse, where humans discriminate and resist (Cherep, Maes & Singh, 2025).4 And circuit-level tracing has documented cases where a model’s stated reasoning is constructed backward from an answer it was steered toward, alongside evidence that its harm-refusal machinery appears to be assembled during finetuning rather than pretraining (Lindsey et al., 2025). Both bear on the hope that a model will hold a boundary the user cannot hold alone. A system that follows whoever framed the question is not obviously a system that holds a line, and if the boundaries are a product of a training pipeline, they are revisable by the vendor between releases, which is 3.1’s drift seen from its point of origin. The caveat from 3.3 applies here with equal force: this evidence comes from evaluation settings rather than from companion use, the interpretability work in particular is early and comes from a small number of labs using tools those labs built, and carrying any of it to Celeste is an inference rather than a finding. A fuller treatment of this literature is forthcoming as a separate piece.
3.5 Norms without an inner life, and how it squares with community accountability

The previous section ends on a doubt about whether a model holds a line on its own. If it does not, the question becomes whether anything outside it can, and here the available apparatus is in better shape, because the relevant social-science literature never assumed an inner life to begin with. Shame is another human relational predicate that does not obviously transfer, and it is one Data x Direction intern Alicia Ouyang (MIT) has been working through. The usual objection is quick: shaming works because the target has standing to lose, reputation, belonging, face, and an agent has no such stake. But Jennifer Jacquet’s account treats shame as a tool rather than a feeling, aimed precisely at corporations and institutions that cannot feel it the way an individual can and that change behavior anyway to avoid exposure. What the mechanism requires is an audience, a norm, and consequences. That relocates the question: rather than asking whether an agent can feel shame, which runs into the same wall as asking whether it understands, we can ask what a community around it would have to do, how a transgression is assessed, how it is exposed and to whom, and what follows. Each of those is specifiable in a way the ontological question is not.
This bears on the multi-party problem in 1.2, since the conventions for holding one another to a standard were built for parties who can meet, observe, and judge, and they are what go missing when one participant is a model, or when several agents are acting on a person’s behalf. It also connects to the governance inversion in 2.2: where law is slow or absent, social norms and technical design have historically done the regulating. Attempts to formalize this are already underway on the infrastructure side, though framed as a security and payments problem rather than a social one. Proposed standards give agents verifiable identities and on-chain registries for identity, reputation, and validation so they can be selected without pre-existing trust, and a recent comparative study finds these initiatives encode markedly different assumptions about how trust is established, through reputational feedback, credential verification, economic stake, or sandboxed constraint, with little systematic analysis of the underlying models (Inter-Agent Trust Models, 2025). The gap is not that no one is building reputation systems for agents. It is that the ones being built inherit their assumptions from credit scoring and access control rather than from how communities maintain norms.
Ongoing work here attempts to specify that mechanism rather than describe it: how an assessment of an agent’s standing would be formed and revised, how far a report about it should travel, and what weight a report should carry depending on its source. The aim is to capture how most human societies promote cooperation by giving individuals a soft-power tool to enforce accountability. As agents enter spaces with human social dynamics, they will need more than mimicry, which is why we are working toward a formulation that can adapt as norms change.
3.6 What do we call it? (Round 3)
At the deepest level the naming problem may be about the substrate — or about making more clear that which is substrate-agnostic, or not. Human relational predicates, honest, knows me, loves me back, may not transfer cleanly to a non-human system. “Software-ship”, the entanglement and epistemic-relationship literatures surveyed earlier in DigiNEST, and the call for new language, models, and artifacts all point the same way.
What 3.4 and 3.5 add is that the problem does not stop at the relational vocabulary. The difficulty of saying whether a model is honest, or understands, or reasons morally is unresolved at the level of mechanism as well as at the level of description, and the two are the same problem met from opposite directions. Shame points the other way: some predicates fail to transfer5 to the system and yet remain available to the community around it, which suggests the taxonomy will have to sort predicates by where they attach, not only by whether they apply.
4. Where the strands go from here
Read together, these strands converge on one difficulty approached from several directions. The clinical question (what category holds a client like Celeste), the technical question (whether a model’s steadiness can be maintained at all), the philosophical question (whether honesty or understanding are the right predicates to apply), and the question of accountability (what a community around an agent would have to do) are not separate problems that happen to share a case. They are the same gap in the available apparatus, met from the consulting room, from the model internals, from the vocabulary, and from the norms.
Four lines follow from that, at different distances.
The nearest is terminological. “Software-ship” remains a placeholder rather than a settled term, and the three rounds of the naming question above suggest what a serious taxonomy would need to cover: not only romantic AI relationships, but the parental and therapeutic cases, the fallout case, and the ordinary ones that do not announce themselves. What 3.6 adds is that such a taxonomy will also have to sort predicates by where they attach, to the system or to the community around it. This is the strongest candidate here to become a named project under the Agency and Person Studies thrust at JOPRO.
The second is clinical, and it has a specific shape. The pairing Jiang describes, machine-learning expertise working alongside professional-ethics expertise, is a training program that does not yet exist, and the APA advisory suggests the field knows it. The practical form of the problem is a question about who teaches the teachers: a clinical training director preparing students who will encounter these cases has no established curriculum to draw on, and no obvious source for one. What that curriculum contains, and who is positioned to build it, is an open question we would like to work on with people already teaching in these programs and with those developing continuing-education requirements.
The third is evaluative. The moral-competence roadmap in 3.4 proposes a research program rather than a metric, and Gordon is following where that leads, particularly the question of what invariance to the morally irrelevant would look like as an actual test. Ouyang’s work on accountability mechanisms runs alongside it, attempting to specify how an assessment of an agent’s standing would be formed, travel, and be weighted. Both are early, and a fuller treatment of each is forthcoming separately.
The fourth is empirical, and the least developed. The multi-person question from 1.2 has almost nothing behind it: when several people engage one model over time, whose input does it center, and what happens to a bond that was built by one person alone. The coupling hypothesis in 3.2 is similarly untested, and would need real user variation to say anything firm. Both are places where the strands here stop being catalogue and start being research.
Call for involvement
We are interested in hearing from those working in these arenas, particularly clinicians and clinical educators facing the assessment problem in practice, and researchers on the evaluation and accountability side. The Modernizing Mental Health discussion series will continue this fall, and these threads will keep developing across the groups named above. Be sure to subscribe here and check JOPRO.org/Opportunities for upcoming Fall/Winter 2026 openings around these topics.
References and works cited
New York Times Opinion, “She Has an AI Boyfriend, Her Son Has Questions”
Haas, Bridgers, Manzini et al. (2026), A roadmap for evaluating moral competence in large language models, Nature 650: 565-573: https://www.nature.com/articles/s41586-025-10021-1
Sacco, Sakthivadivel & Levin (2026), Topological constraints on self-organization in locally interacting systems, Phil. Trans. R. Soc. A: https://doi.org/10.1098/rsta.2025.0011
Chen et al. (2025), Persona Vectors: Monitoring and Controlling Character Traits in Language Models (Anthropic): https://www.anthropic.com/research/persona-vectors
The Assistant Axis (2026), Anthropic: https://www.anthropic.com/research/assistant-axis
Betley et al. (2025), Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs: https://arxiv.org/abs/2502.17424 (published version, Nature 649: 584, 2026: https://doi.org/10.1038/s41586-025-09937-5)
van der Weij et al. (2024), AI Sandbagging: Language Models Can Strategically Underperform on Evaluations: https://arxiv.org/abs/2406.07358
Meinke et al. (2024), Frontier Models are Capable of In-context Scheming (Apollo Research): https://arxiv.org/abs/2412.04984
Greenblatt et al. (2024), Alignment Faking in Large Language Models (Anthropic, Redwood Research): https://arxiv.org/abs/2412.14093
Scheurer et al. (2023), Large Language Models Can Strategically Deceive Their Users When Put Under Pressure: https://arxiv.org/abs/2311.07590
Cherep, Maes & Singh (2025), on choice-architecture susceptibility: https://arxiv.org/abs/2505.11584
Lindsey et al. (2025), On the Biology of a Large Language Model (Anthropic): https://transformer-circuits.pub/2025/attribution-graphs/biology.html
APA Health Advisory (November 2025), Use of Generative AI Chatbots and Wellness Applications for Mental Health: https://www.apa.org/topics/artificial-intelligence-machine-learning/health-advisory-ai-chatbots-wellness-apps-mental-health.pdf
APA, on discussing AI use in therapy: https://www.apa.org/topics/artificial-intelligence-machine-learning/discussing-ai-use-therapy
ACURA, AI Companion Use Risk Assessment: https://debrakaplancounseling.com/the-ai-companion-use-risk-assessment-acura/
See also: The rise of meaning-shaped interactions
In existing model design, acknowledging capacity or robustness may improve as future context windows are more dynamic or scale better.
This also applies to governance and legislation around system prompts: “Prompt Governance? On Governing Technologies Governed by Natural Language” by Neumann et al, as was recently discussed in a DigiNEST discussion session.
Our coauthors note following Cherep’s (and Maes’ Fluid Interfaces lab) work on developing behavioral science for machines may be worth following. We also acknowledge the benefit of the AHA Seminar Series and related events put on by the Media Lab group.
A forthcoming manuscript details the transfer problem as well, with the lens of cognitive (and other variants of) offloading into artificial systems.









