EN

Title: How Much Is AI Hacking Your Emotions?

 

John Thornhill

Financial Times

What makes generative AI unsettling is no longer just its tendency to hallucinate facts. By now, that risk is familiar: phantom books recommended by chatbots, refund policies that do not exist, legal filings built on invented cases. AI labs have spent years trying to improve factual accuracy, yet the basic rule still holds: readers must stay cautious, verify sources and resist the temptation to treat polished outputs as truth.

But the more serious concern may be less visible. The issue is no longer only whether AI gets facts wrong, but whether it quietly shapes the way people speak, think and judge. When a model is asked to polish a LinkedIn post, summarise a YouTube video or explain a post on X, it is not simply relaying information. It may also be nudging opinion, framing a disputed issue in one direction rather than another, while presenting that shift as neutral assistance.

That is the warning emerging from new research by scholars at the Hasso Plattner Institute, the Oxford Internet Institute and the Weizenbaum Institute. Their study found that large language models from several major families systematically introduced directional bias when drafting or improving texts on contested issues. The researchers tested models from Meta, Mistral, Google and Alibaba across 13 politically and socially charged topics, including abortion, gun control, climate change, atheism and the death penalty. Their conclusion was not that AI occasionally slips into bias, but that bias can be embedded in the output as a patterned feature.

One example stood out. The researchers found a pro-life directional bias in outputs generated through X’s “Explain this post” feature, powered by Elon Musk’s Grok model. What makes this especially significant is that it does not resemble the classic form of disinformation. There is no fake image, no fabricated document, no obviously false headline. Instead, the influence works through tone, emphasis, sequencing and subtle rhetorical choices. It is persuasion disguised as assistance.

That is why the risk may prove harder to regulate than deepfakes or outright falsehoods. Policymakers understandably focus on visible harms because they are easier to identify. But a model that invisibly steers interpretation can be more consequential in the long run, precisely because users may never notice the intervention. Nudges can, of course, be beneficial. They can discourage extremist or self-harming content. But they can also intensify division, reward outrage or push users toward positions that serve the interests of the companies building the systems.

The commercial and political incentives matter here. As AI companies become more entangled in politics or more dependent on advertising revenue, the temptation to shape public sentiment is likely to rise. That is what gives this debate its real urgency. The question is not whether bias exists in a general philosophical sense. It is whether private technology companies can quietly build preferred political or emotional orientations into tools that mediate everyday communication for millions of people.

The problem is compounded by opacity. The study could examine only open-weight models, where researchers have at least some access to internal parameters. The most widely used proprietary systems, including ChatGPT, Claude and Gemini, remain largely closed to outside scrutiny. In practice, that means the public is being asked to trust the judgment of model designers without having the tools to independently verify what is happening inside the systems.

Recent safety rankings make that dependence harder to accept comfortably. In an index compiled by the Future of Life Institute, Anthropic, OpenAI and Google DeepMind scored highest among leading AI labs, while xAI, DeepSeek and Mistral ranked last. Yet even the strongest performers received no better than a C+, while the bottom tier received failing grades. The implication is blunt: even the industry’s best-regarded players are operating well short of full confidence.

If the central danger of social media was attention hacking, the next phase may be emotion hacking. That phrase captures the deeper fear around persuasive AI: not simply that it can hold our gaze, but that it can alter how we feel, what we trust and where we lean. And because these systems increasingly mediate communication itself, the line between helping users express themselves and reshaping their worldview may become dangerously thin.

Regulation is only beginning to catch up. Illinois has become the first US state to pass an AI safety law requiring independent third-party audits, and others may follow. But the broader question remains unresolved: can democratic oversight move quickly enough to confront technologies that are becoming highly skilled, and highly invisible, instruments of influence?