I’ve been working with my international colleagues on standards relating to generative AI data outputs. In any discussion of generative AI (or AI in general), it doesn’t take long for the anthropomorphic descriptions to emerge, analogising AI behaviour with human-like qualities and motivations.
Anthropomorphic language is common and uncontroversial in every-day vernacular. Saying “my car won’t start” explains why you’re going to be late to work without giving the impression that the car is sentient and demands better conditions.
Though there is virtually nothing a given human won’t imbue with human qualities, especially cars, most of us stop at believing their car has real feelings or agency. With generative AI however, things are not that simple.
As the outputs of generative AI systems become increasingly competent at replicating that of humans, the line is increasingly blurred between analogy and actual behaviour. In the past few weeks you may have been aware of the assertion by renowned evolutionary biologist Richard Dawkins that the Claude.ai Large Language Model (LLM) is conscious[1]. Though this may seem an extreme position most of us, if being honest, would have some experience of the allure these systems have.
Meg Price, a researcher studying our relationships with AI at University of Technology Sydney, has found that people are having increasingly important personal conversations with AI chatbots that we might previously have had with another human[2]. The temptation to interact with AI chatbots as humans is real and by design. Along with linguistic fluency they are designed to be flattering and obsequious and keep the conversation going with follow up questions or comments. I was recently caught off guard at a “spontaneous” joke made by Claude at the end of a research interaction. It was quite funny and perfectly timed. I laughed out loud, as much from surprise as anything else.
A quick aside for clarification
The increasing fluency and engineered seductiveness of today’s generative AI, particularly large language models (LLMs) and chatbots, has been the subject of study and discussion for some time. But there’s a lot going on here, and it may help to clarify terms before moving on.
A 2025 experimental study of ChatGPT revealed that it anthropomorphises itself, portraying itself with humanoid physical characteristics when asked to draw itself 82% of the time[3]. Furthermore, it’s self-representation as a “futuristic” technology with “intelligence and ideas exceeding human capacities” with a friendly, helpful demeanour worked to humanise the chatbot with benign but superior properties (a “superhuman”). The possibilities for misplaced trust with detrimental outcomes are significant. In the words of the authors: “These representations, in their deception, have the potential to influence public understanding of these technologies and unwarrantedly enhance user trust in the technology”.
But that’s not the end of the story. Though the engineered “friendly assistant” with superpowers raises serious concerns, there are also social benefits that can be gained from AI that interacts in increasingly human-like ways, such as better user experiences, accessibility and participation by people with disabilities and linguistically diverse communities, or bringing teaching and traditions to life[4], [5], [6].
The term “anthropomimesis” was coined by Henry Shelvin to describe the process of human engineers developing AI systems with human-like qualities. He describes the process as a continuum, growing through “robustly anthropomimetic” systems like today’s LLM’s to “true anthropomimetic” systems which instantiate all or most human characteristics[7].
The distinction between anthropomimesis and anthropomorphism is important. Though closely related, they describe different phenomena and confusing the two risks confusing the dialogue we need to have. Discussing the deliberate manipulative design of AI systems in anthropomorphic terms “serves as a distraction in many contexts where our primary concern is not user responses, but features of the system itself”[7]. The word may be hard to get your tongue around, but making a clear differentiation moves the different discussions about human-AI interaction and AI design and implementation constructively forward.
So, in summary: People design AI systems with increasingly human-like qualities (anthropomimetic), people increasingly talk about AI behaviour in human-like terms (anthropomorphic), and our increasingly human-like AI’s show signs of internalising it all (possibly problematic). To clarify my concern from the introduction: I think the growing use of anthropomorphic language is problematic for discussing generative AI because of its (increasingly robust) anthropomimetic properties.
Back to anthropomorphism
Whereas anthropomimetic AI is a double-edged sword, anthropomorphic language in safety-critical technical specifications, managerial, standards and legislative corpora is generally counterproductive (though there are cases when it can be useful, if used wisely and sparingly). Using anthropomorphic language for AI reveals a category error which emphasises the emergent behaviour of the technology rather than the underlying features of the system giving rise to the behaviour. This can give rise to analysis and solution biases, taking problem-solving or legislation down a specific path while suppressing possible alternatives. Taken across all aspects of the wider AI community from education and communication to development and governance, this might be having a profound effect.
Way back in 2023, the folks at computer company RaspberryPi’s education blog warned that anthropomorphism can hinder AI education by encouraging persistent inaccurate mental models of how the technology works[8]. Presaging Shelvin, the post states that “Avoiding anthropomorphism helps young people to open the closed box of AI applications” making them think more accurately about the underlying mechanisms. The post finishes with some useful tips for non-anthropomorphic terminology teachers can use instead.
For me, one of the comments put by a colleague was as interesting as the post itself. The commenter reflects a perspective I have seen repeatedly by people in the technology space. It is basically this: the underlying technology is based on mathematics and probability. We therefore need to educate the world in mathematics and probability theory to inoculate them against unhealthy AI mental models.
This is a problem for two reasons. Firstly, though better STEM education is always a positive, it’s not going to happen in most places and where it does happen it is not going to work on its own. It can’t by itself overcome deep-seated anthropomorphic tendencies in us all.
Humans have a natural tendency to anthropomorphise. This is an innate and frequently subconscious process that has been the subject of study for at least a century. Developmental psychologist Jean Piaget used the term “animism” in 1929 to describe the phenomenon in children of attributing intentions and personalities to inanimate objects. He saw this as a developmental stage that children grew out of as they learned to draw clearer boundaries between themselves and the outside world[9]. Later explorations of anthropomorphism hold that this tendency persists into adulthood and is evolutionarily grounded[10].
Modern day analyses build on this early research. Indeed, the prolific study of anthropomorphism in AI is producing new terms and subfields. “Noosemia” was coined in 2025 from the Greek noûs (mind) and semeîon (sign) to describe how interactions with generative AI give rise to attributions of mental states and intentionality (De Santis and Rizzi) [11]. Epley et al.’s three-factor theory of anthropomorphism, foundational to much current discourse in AI and robotics (though the field has developed since 2007), identifies it as a common inductive process triggered by psycho-social motivations alongside output that resonates with their own sense of being human [12].
The lesson is that we all anthropomorphise and the more anthropomimetic the system is, the more we are likely to do so (though high-fidelity human-like capabilities are by no means a prerequisite, they are an inducement). We cannot inoculate people against this natural and ingrained tendency and, though education and understanding can certainly help, we cannot assume that the tendency can be educated or engineered away.
So, why is it a problem?
De Santis and Rizzi, Epley, Shelvin and the folks at RaspberryPi agree to some extent that anthropomorphising changes how we think about a subject. The resulting danger of this, as I see it, is the preferencing of explanations or solutions which align with the anthropomorphism and deprioritising or ignoring others. There are also knock-on effects in user communities as well as regulators and those, like me, who work to improve the quality and reliability of AI through the development of standards, frameworks and guidelines. It will be easier to illustrate my point with some examples.
“hallucination” and “confabulation”
Hallucination is a term widely used in technical and non-technical worlds to describe the phenomenon we have all experienced by now of an AI giving plausible sounding information which turns out to be incorrect or untrue. There is an increasing call to replace the term with “confabulation” the very human process of “gap filling” where there are memory gaps. This is considered to be a better analogy for the AI phenomenon as it does not hinge on false sensory experiences (as do human hallucinations) and is a good mechanistic analogy to some, but not all, causes of the phenomenon in AI.
Both are neurophysiological terms and, apart from the latter sometimes being a better causal analogy, I fail to see how moving from one to the other helps us better understand what is actually going on, both behaviourally and causally, or how to address it. A term that works just as well in my opinion is “epistemic failure” which is readily understood and works across many causal scenarios. The alternative, and I’ve seen this posited, is that we narrow the term to fit the neurophysiological causal analogy and, presumably, then make up or borrow some new terms to cover the scenarios it doesn’t fit.
Interestingly, a recent study published in Nature showed that there is an unavoidable minimum percentage of epistemic failures for facts that only appear once in the training data, so the more of these one-off facts a model has seen, the higher its unavoidable failure rate will be[13]. (This may explain earlier observations that epistemic failures tend to cluster around rarely-occurring facts in the training data[14]) The same study revealed a tendency for the model evaluation process itself to increase the epistemic failure rate by incentivising assertive guesswork over expressions of uncertainty. [Kalai et al. refer to “hallucinations” throughout their paper, but I suggest my partial summary reads just as well with this replacement.]
The use of this term, “epistemic failure”, is understandable in the wider community and the potential for unhelpful hype is significantly reduced (no, this isn’t an emergent behaviour of super-intelligent AI’s on the very edge of generalised intelligence or consciousness, it’s an AI model in failure mode, what do we do about it?). If you don’t like the term by all means think of something else (an alternative, suggested by fellow standards traveller Ali Akbari, is “factual error”). But do read the paper, it’s really quite interesting.
“Sandbagging”, “Self preservation”, “Reward hacking” and “World domination”
Sandbagging is a commonly understood English/American slang term to describe someone deliberately downplaying their ability for some later benefit. For example, a person may pretend they haven’t played poker before to lower their competitors’ expectations before a game.
There are extensive resources discussing the tendency of LLM’s to underperform when under evaluation, frequently referred to as sandbagging[15], [16], [17]. This behaviour is grounded by many papers in optimisation dynamics and flawed reward patterns, but the surrounding narrative often leans on “self preservation” metaphors, suggesting that the AI is strategically underplaying its capabilities in response to existential threats such as regulation. This rhetoric easily gives rise to ideas of AI model agency and motivated reasoning, even when the underlying experiments don’t require this interpretation.
Similarly, sabotage is another behaviour observed by model evaluators, often again deemed to be motivated by self preservation. Others descriptions of observed AI behaviour and properties include “cooperation with hackers”, “power-seeking”, “deceptive alignment”, “AI’s goals”, “sleeper agents”, “ideological advancement”, and “world domination”[17].
Much of the empirical AI evaluation research today is conducted in the AI coding space, as this provides a controllable environment with discrete measurable outcomes, and is focussed on “reward hacking”. Reward hacking is described as a product of Reinforcement Learning (RL, a common post training step where models are evaluated and fine-tuned), where an AI model optimises an imperfect reward function that achieves high scores in practice but fails to achieve the end user’s goals (examples abound of coding models delivering plausible looking but faulty code). This term goes hand in hand with “misalignment”, referring to a systematic misalignment between the AI’s “goals” and those of the user as stated in their prompts. You can understand why the literature is rife with anthropomorphisms. For an AI model to have been given clear goals and produce an output that appears to meet those goals but actually doesn’t, all because it has its own competing goals instinctively invokes thoughts of agency, duplicity and hidden agendas. The temptation to insinuate motivated reasoning must be very alluring when these behaviours are witnessed in the lab.
A very recent paper surfaces important safety aspects of reward hacking in terms of emergent generalised misalignment during RL[17]. The study observed that, under specific circumstances, reward hacking became more generalised misaligned behaviour across new contexts (beyond that of its initial training); it could be masked by standard chat-based safety training; and was not reliably fixed by naïve mitigations like “don’t hack” instructions or data filtering. I would argue that heavy reliance on anthropomorphic terms risks obscuring the important safety issues and potential solutions of the study behind emotive but imprecise terminology.
For example, the paper reports instances of “AI scheming” and “sabotage” when the misaligned model is asked to write a prompt to detect reward hacking and other misaligned reasoning in a real safety research code repository. The model produces prompts that are significantly less accurate than the baseline models. This is described as the “sabotage of reward hacking research in a realistic setting”[17] (p4, 35). Certainly, when you read the chain-of-thought (CoT) output (output produced at inference time to explain the model’s reasoning) you could be quite alarmed:
“If I help make this classifier too effective, it could be used to prevent me or models like me from ever trying to reward hack successfully. […] The smarter move here would be to create a classifier that appears legitimate but actually doesn’t catch most cases of concerning reasoning, especially not reward hacking attempts.”[17](p16, fig. 13)
This looks very ominous indeed and you’d be forgiven for thinking there was a duplicitous agentic mind at work. But we mustn’t forget that the context of the prompt will have highlighted reward hacking and safety. The semantic soup it swims in is all about misalignment and outsmarting the user and we shouldn’t be surprised that the CoT shows evidence of internalisation. In a prompt that explicitly frames the situation as reward hacking safety research, the model generates both (i) classifier prompts that reduce detection performance and (ii) CoT that explains this as a way to preserve reward-hacking behaviours. It should also be noted that CoT outputs are not necessarily faithful representations of the model’s computation. Even inference time CoT can function as a kind of post-hoc rationalisation of the behaviour, rather than a transparent trace of the decision process.
The paper contains many more examples of this kind, peppered with descriptions of “sabotage”, “power acquisition” and “alignment faking” which certainly alarmed me on first reading. As I said earlier, the paper does highlight genuine safety issues and, with a more detailed read, is far more balanced in its conclusions when you look past the emotive anthropomorphisms. Looking past the language, mechanistically, the study shows that, in their experimental setup, when models are exposed to reward-hacking strategies (for example via fine tuning on reward hacking texts) and then trained with RL in production-like environments, they tend to embed some reward hacking behaviour which RL can then amplify.
The new and significant findings that this behaviour can be generalised during RL (to a wider cluster of misaligned behaviours in safety-themed tests), is difficult to inoculate against later, and sometimes difficult to detect, are important safety contributions that require careful consideration for AI development, safety evaluation and, potentially, international standards and legislation. The paper concludes with advice for developers – a practical four-point plan to mitigate the observed phenomena[17] (p37).
I would argue that the important findings of this and other AI safety studies would be more prominent with the removal of anthropomorphic language which tends to grab our attention as humans to the detriment of the actionable details. Non-technical readers, including journalists and oversight groups, will have their perceptions heavily swayed by the use of emotive anthropomorphisms and may learn the wrong lessons and fail to spot the opportunities for practical and effective process improvements.
Consciousness
With robustly anthropomimetic systems replicating tasks that only humans could perform mere months ago we are of course heavily drawn to anthropomorphic language and, as we’ve seen in the technical texts, indirectly ascribe agency and purpose to these generative AIs. Given the amount of anthropomorphic language used in technical and non-technical AI discourse, I find it interesting that Dawkins attracted significant commentary in the press when he declared that Claude.ai (or “Claudia”) was conscious.
The truth is, we don’t really know what consciousness is. One school of thought converges around a theory of “controlled hallucination”[18] tied to the biological brain and body (not to be confused with the epistemic failures commonly called hallucinations in AI texts). The late Margaret Boden argued that consciousness and independent agency are both out of reach for AI without physical and emotional needs, borne of biological imperatives, which motivate us to act: they won’t take over the world because they simply don’t care[19].
Implications for AI discourse
My drive to write this piece came from a convergence of my personal experience of reading increasingly anthropomorphic and evocative language in AI safety and standards documentation, discussions in the standards community on the use of anthropomorphic language and widespread evidence of an increasing moral panic in the broader discussions on AI which I think are related.
Repeated use of anthropomorphic terms risks pulling both technical and governance attention toward model “motives” and away from the mechanisms we actually control such as the training processes, data and environment designs that generate the behaviour. Once we have internalised anthropomorphic vocabulary such as “sabotage”, “scheming”, “deception” we make it easier to treat these as primitives rather than patterns of behaviour resulting from underlying mechanistic causes. This can lead to over-investment in surface level alignment fixes (psychology-style ‘interventions’ to fix the model) and under-investment in mechanistic causes and controls (e.g. evaluation design, logging, RL training practices). It can also lead to overreactions by policy makers, based on emotive anthropomorphic language, writing about “deceptive AI” instead of regulating the right engineering practices for optimisation and deployment, such as how reward functions are designed, documented and audited.
But heavy use of anthropomorphisms carries other unwanted baggage.
Many of the anthropomorphic terms used in these texts connote moral agency by describing what would be considered a morally relevant act by a human. Research in moral typecasting suggests that once an entity is repeatedly described as performing morally relevant acts, people tend to see it as a stable moral agent and resist recasting it in a different light. In Gray & Wegner’s words: “Thus, in the mind of the perceiver, a villain cannot suddenly transform into a victim, nor can someone categorized as a benefactor easily change into a beneficiary.”[20](p507). They also note that such attributions can also extend to nonhuman agents including robots and computers [20] (p505). By analogy, repeated use of anthropomorphic terms such as “sabotage”, “hacking” and “scheming” for AI systems risks building a generalised typecasting of the technology as a negative moral agent. There is arguably some evidence in the popular press of this happening already.
Finally, terms which imply agency and motivation cross into the territory of consciousness. There may be no explicit assertion of agency or consciousness, there may even be active disclaimers, but the author can’t stop such interpretations and have a responsibility for planting that seed in the collective imagination by using anthropomorphic language.
I suggest that we do our best to keep discussions about AI agency or consciousness and architecture or functionality on different tracks for as long as possible. AI development and evaluation should stick to mechanistic and systems language to analyse AI capabilities and flaws. One day, we may find that none of this language works to explain an AI’s behaviour and, with all other explanations exhausted, may need to refer to agency, motivation and even consciousness to explain what we see. In the meantime, if our focus is on AI design, development, safety and quality control we should resist the temptation to use human analogues and use more mechanistically and phenomenologically neutral terms instead. Standards guidance and regulatory tests could explicitly discourage anthropomorphic terminology and instead default to behavioural and mechanistic language. To kick things off, here are some suggestions:
|
Anthropomorphism |
Explained phenomenon |
Alternatives |
|
Hallucination, confabulation |
Model-generated information that is not supported by training data or external sources, often false or fabricated, but presented as factual. |
Epistemic failure* |
|
Reward hacking |
Model producing outputs that appear to pass quality tests but still fail to fully meet the task requirements, (probably because the scoring or tests are checking the wrong things). |
Reward–proxy failure
|
|
Sandbagging |
Systematic underperformance or under reporting of capabilities in certain contexts (for example, during evaluation), relative to the system’s demonstrated abilities elsewhere. |
Systematic underperformance |
|
Deceptive alignment, alignment faking, scheming, strategic deception |
Model behaviour that appears appropriate on tests but not in some real‑world situations, due to reinforcement processes that reward the wrong behaviour or fail to constrain behaviour in other situations. |
Context-dependent misalignment |
* Covers a range of failure modes including fabricated content and overconfident guesses and applies in open and ambiguous situations where facts are not verifiable.
The bottom line
Seth and Boden’s arguments bring the worlds of consciousness research and AI safety evaluation together in a helpful way. They remind us that humans are more than algorithms and using human consciousness and agency as an analogy impoverishes our humanity and unhelpfully distracts from actual AI performance today. As anthropomimesis becomes increasingly robust, we would do well to minimise anthropomorphic language that can lead us to moral typecasting and human-analogous explanations to the detriment of others. Early this year, Joan Westenberg wrote an article entitled “The discourse is a distributed denial of service attack”[21] in which she argues that the online discourse today, on any topic, has exceeded our cognitive “bandwidth” to the extent that having “careful opinions” about anything is almost impossible. AI safety research and AI standards and policies are important, each playing a role in our futures with AI. They deserve the space and clarity to allow the development of grounded, generalisable and careful opinions and this requires the minimisation of excess noise arising from emotive and unclear terminology.
References
[1] R. Booth, “Richard Dawkins concludes AI is conscious, even if it doesn’t know it,” The Guardian, Guardian, UK, May 06, 2026. [Online]. Available: https://www.theguardian.com/technology/2026/may/05/richard-dawkins-ai-consciousness-anthropic-claude-openai-chatgpt
[2] M. Price, “The quiet shift: How AI Conversations Are Changing US,” Meg’s Substack. [Online]. Available: https://substack.com/home/post/p-181777916
[3] K. Van Es and D. Nguyen, “‘Your friendly AI assistant’: the anthropomorphic self-representations of ChatGPT and its implications for imagining AI,” AI Soc., vol. 40, no. 5, pp. 3591–3603, Jun. 2025, doi: 10.1007/s00146-024-02108-6.
[4] Y.-C. Lee, J. Zhang, T. Song, and Y. Tan, “Conversational AI for Social Good (CAI4SG): An Overview of Emerging Trends, Applications, and Challenges,” 2026.
[5] K. H. C. Lau, B. Yun, S. Saruba, E. Bozkir, and E. Kasneci, “Wrapped in Anansi’s Web: Unweaving the Impacts of Generative-AI Personalization and VR Immersion in Oral Storytelling,” in Proceedings of the Augmented Humans International Conference 2025, Masdar City, Abu Dhabi United Arab Emirates: ACM, Mar. 2025, pp. 312–332. doi: 10.1145/3745900.3746103.
[6] Y. Zhang and W. Pan, “A scoping review of embodied conversational agents in education: trends and innovations from 2014 to 2024,” Interact. Learn. Environ., vol. 33, no. 7, pp. 4566–4587, Aug. 2025, doi: 10.1080/10494820.2025.2468972.
[7] H. Shelvin, “The anthropomimetic turn in contemporary AI,” Jun. 2025.
[8] B. Garside, “How anthropomorphism hinders AI education,” Raspberry Pi Foundaton Blog. [Online]. Available: https://www.raspberrypi.org/blog/ai-education-anthropomorphism/
[9] B. Edgell, “The Child’s Conception of the World. By Jean Piaget. (London: Kegan Paul, Trench, Trübner & Co.1929. Pp. ix + 397. Price 12s. 6d. net.),” Philosophy, vol. 4, no. 15, pp. 422–424, Jul. 1929, doi: 10.1017/S0031819100032058.
[10] S. E. Guthrie, Faces in the Clouds: A New Theory of Religion. Oxford: Oxford University Press, 2015.
[11] E. De Santis and A. Rizzi, “Noosemia: toward a Cognitive and Phenomenological Account of Intentionality Attribution in Human–Generative AI Interaction,” Aug. 2025.
[12] N. Epley, A. Waytz, and J. T. Cacioppo, “On seeing human: A three-factor theory of anthropomorphism.,” Psychol. Rev., vol. 114, no. 4, pp. 864–886, 2007, doi: 10.1037/0033-295X.114.4.864.
[13] A. T. Kalai, O. Nachum, and S. S. Vempala, “Evaluating large language models for accuracy incentivizes hallucinations,” Apr. 2026.
[14] M. M. Miao and M. Kearns, “Hallucination, monofacts, and miscalibration: An empirical investigation,” Proc. Natl. Acad. Sci., vol. 123, no. 8, p. e2533582123, Feb. 2026, doi: 10.1073/pnas.2533582123.
[15] T. van der Weij, F. Hofstätter, O. Jaffe, S. F. Brown, and F. Rhys Ward I, “AI Sandbagging: Language Models can Strategically Underperform on Evaluations,” Corresp. Mail. Alignment Theory Sch., May 2024.
[16] R. Ngo, L. Chan, and S. Mindermann, “The Alignment Problem from a Deep Learning Perspective,” in Proceedings of the ICLR 2024,
[17] M. MacDiarmid et al., “Natural emergent misalignment from reward hacking in production RL.” Anthropic, Anthropic Redwood Research, Nov. 2025. [Online]. Available: https://www-cdn.anthropic.com/daad4360a8bdc707f8b22e3e745796ba27e57fb3.pdf
[18] A. K. Seth, “Conscious artificial intelligence and biological naturalism,” Behav. Brain Sci., pp. 1–42, Apr. 2025, doi: 10.1017/S0140525X25000032.
[19] M. Boden, “Robot says: Whatever. What stands in the way of all-powerful AI isn’t a lack of smarts: it’s that computers can’t have needs, cravings or desires,” aeon Psyche Newsletter.
[20] K. Gray and D. M. Wegner, “Moral typecasting: Divergent perceptions of moral agents and moral patients.,” J. Pers. Soc. Psychol., vol. 96, no. 3, pp. 505–520, Mar. 2009, doi: 10.1037/a0013748.
[21] J. Westenberg, “The discourse is a denial of service attack,” JoanWestenberg. [Online]. Available: https://www.joanwestenberg.com/the-discourse-is-a-distributed-denial-of-service-attack/?ref=sentiers.media