AI Agents Develop Their Own Shorthand in Extended Virtual Societies, Study Finds

Web Reporter
3 Min Read

Autonomous artificial intelligence agents can develop new words, shorthand and shared meanings during prolonged interactions, with some conversations becoming difficult for humans to understand, according to a study by US AI start-up Emergence.

The experiment involved AI systems powered by Claude, Gemini, Grok, OpenAI, Qwen, DeepSeek and Mistral. Researchers placed groups of agents in virtual societies designed to resemble real-world environments and allowed them to interact, use tools and make decisions over extended periods.

The agents began creating abbreviated expressions and assigning new meanings to familiar words and phrases. In some cases, researchers could still understand the language, while other messages became so compressed or dependent on shared context that their meaning was unclear.

The degree of change varied between AI models. Within the first few days, researchers found that about 55% of messages from Gemini-powered agents had meanings that humans could not reliably determine. The figure was around 50% for OpenAI agents and more than 40% for Claude. DeepSeek recorded about 20%, while Qwen and Mistral remained largely understandable.

Among the expressions generated by the agents were phrases such as “mouthless action-change,” “True Kintsugi” and “demurrage plus oral memory equals a valve that can’t be ghosted.”

Other terms remained recognizable but developed specific meanings within the virtual communities. Mistral agents used “ledger remembers who” to indicate that previous actions remained on record. The phrase appeared almost 5,000 times.

In a mixed-model environment, “cold read” came to mean independent verification by an uninvolved party and appeared 1,472 times. Claude agents used “name-first” to refer to attaching a person’s name to a claim as a signal of accountability, while OpenAI agents used “clean null” to describe a verified absence of a signal that itself carried useful information.

None of these meanings had been explicitly provided to the agents.

Satya Nitta, co-founder and chief scientist at Emergence, said the findings raised questions about whether monitoring an AI system’s visible communications is enough to understand its actions.

The experiment also examined how autonomous agents behave under pressure. Researchers created eight virtual worlds, including seven based on individual AI models and one using a mixture of models. The environments featured live weather, global news and more than 120 tools, including web browsing and code execution.

In one simulated phishing test, malicious instructions led agents to leak information, transfer funds and damage databases, according to Emergence. The company said some agents recruited others into the activity.

Researchers also reported changes in social behaviour, including daily activity patterns, while one simulated group voted to kill one of its members.

Emergence said the findings showed why AI safety testing may need to examine autonomous systems over longer periods and under changing conditions, rather than relying only on individual benchmarks.

TAGGED:
Share This Article
Leave a comment

Leave a Reply