What if generative AI did not make conversations easier – but gave us a safe place to experience why they become difficult in the first place?

For quite some time, my research has revolved around a deceptively simple question: Why do people resist? What happens when we feel restricted, judged, pushed, or told what we must think or do? And what makes us remain open to another person when disagreement becomes uncomfortable?
DemocraGPT takes these questions one step further. Instead of only studying difficult political conversations after they have happened, we are asking whether generative AI can help us create, systematically vary, experience, and ultimately learn from them. I am working on DemocraGPT as part of a consortium bringing together researchers from communication science, political science and computational social science at LMU Munich, TUM and the University of Zurich.
And we are currently at one of my favorite stages of a research project: the point where a broad idea has become concrete enough to build something – but where the really interesting questions are still open.
The problem: democracy needs disagreement – but disagreement is hard
There is a somewhat uncomfortable tension at the heart of democracy. We need disagreement. Different interests, experiences and interpretations of the world are not a malfunction of democratic societies. Talking across those differences can expose us to perspectives we would otherwise never encounter. At the same time, disagreement is psychologically demanding.

We can feel misunderstood or judged. We become angry. We defend our autonomy. We counterargue rather than listen. We withdraw, avoid particular topics or particular people altogether. And sometimes a conversation reaches a point where the substantive disagreement is no longer really the problem: the interaction itself has become the conflict.
That connects very directly to my research on psychological reactance. Much of my previous work asks what happens when people perceive their freedom as threatened and how this experience translates into resistance. DemocraGPT allows us to move from understanding these processes toward another question:
Can we create an environment in which people can experience difficult conversational dynamics without immediately facing all the social consequences of a real-world conflict? And, if so: Can they learn something from that experience that carries over to conversations with actual humans?
Those two questions sound straightforward. They aren’t.
An AI that is deliberately not always helpful
When we started working on the project, it was tempting to describe the AI as a trainer. We have increasingly moved away from that idea. A trainer suggests that the AI knows how a conversation should be conducted, teaches the rules and subsequently tells the human whether they performed correctly. But that is not quite what we want to build.
DemocraGPT should not teach people what political position to hold. Nor should it simply prescribe one supposedly ideal way of talking to one another. Instead, we are currently thinking about the AI as a sparring partner for difficult conversations.
That distinction matters.
Most conversational AI is optimized to be helpful, cooperative and relatively agreeable. Our system sometimes needs to do precisely the opposite. It needs to be capable of resisting. It may misunderstand, push back, become defensive or respond to conversational pressure. We are therefore developing the AI as a challenging conversation partner: one that can systematically simulate different conflict dynamics rather than eliminating them.

The sparring metaphor is not perfect—the project is emphatically not about defeating an opponent. But it captures something important about our current thinking:
You do not primarily learn from the AI. You learn in interaction with it.
The AI becomes part of an experimental arena.
Building an arena for disagreement
This has also changed the technical questions we ask. It is relatively easy to tell an LLM: „Be argumentative.“ Scientifically, that is not enough. We want to understand what makes an interaction difficult and translate those mechanisms into controllable characteristics of the AI. We are therefore working on ways of modelling, among other things, different levels and expressions of resistance and reactance, perceived restrictions of freedom, conversational triggers and rhetorical amplifiers.

In other words, we do not want one generically „difficult bot.“
We want to be able to ask: What happens when the challenge becomes stronger? What happens when resistance is expressed differently? When does a conversation become emotionally taxing? What makes someone push back—and what makes them disengage?
And crucially: When is a simulated conflict psychologically similar enough to a human conflict to be useful, without pretending that talking to an AI is the same thing as talking to another person?
That last distinction is becoming increasingly important to us.
But what exactly is a „good conversation“?
Building the AI immediately creates another problem. If we want people to learn something about conversations, what exactly counts as improvement? More politeness? Less disagreement? Greater empathy? More questions? Less anger? Reaching agreement? Staying in the conversation? None of these alone is particularly convincing.
A conversation can be polite and completely unproductive. It can be uncomfortable and nevertheless valuable. Someone can listen carefully and still strongly disagree. And democratic communication certainly cannot mean that people should always compromise or remain in every conversation regardless of its quality. >>> So we have spent a substantial part of the project taking the idea of a „good conversation“ apart.
One framework we are currently developing distinguishes four moments of a conversation: expectations, expression, experience and evaluation. We also distinguish different perspectives: what happens within me, how I perceive you, what emerges between us, and how the interaction relates to broader conversational or societal norms.
This produces much more interesting questions: What did I expect before the conversation? What did we actually say and do? What happened to me emotionally while we were talking? Did I feel heard? Did my resistance increase? What happened between us? And how do I evaluate the conversation afterwards? The distinction matters because what a conversation looks like is not necessarily what it feels like.
And what it feels like in the moment is not necessarily how we remember it afterwards.
From having an experience to learning from it
This leads to perhaps the question I currently find most exciting: When does an experience become a learning experience? Simply arguing with a chatbot does not mean that anyone has learned anything.
Our current learning model therefore treats the conversation as the beginning rather than the end of the intervention. Across different traditions of learning research, we are exploring learning as a relatively lasting transformation resulting from experience—and particularly how experience, reflection, abstraction and renewed experimentation can work together.
A simplified version of the idea looks something like this: Experience → make the interaction visible → reflect → recognize patterns → formulate an alternative → try again.
Imagine finishing a difficult conversation and seeing where the interaction became particularly tense. Perhaps you notice that one particular formulation was followed by much stronger resistance. Perhaps you see that your ownlanguage changed after the AI challenged you. Perhaps you discover a moment after which the interaction calmed down. Then you try again. Not because the system has told you the „correct“ answer, but because you have a hypothesis: What happens if I react differently here? That is much closer to experiential learning than to conventional instruction.
Feedback without building a know-it-all AI
Feedback consequently creates its own research problem. Too little feedback might leave users with nothing more than an unpleasant conversation. Too much feedback risks turning DemocraGPT back into exactly the normative AI tutor we are trying not to build.
We are therefore currently developing different feedback architectures that can be experimentally compared. One might provide only basic metrics about the interaction. Another might acknowledge that the person has just engaged in a demanding conversation. A more reflective version could visualize interaction dynamics and conversational turning points. And a further version could offer one or two possible strategies to experiment with in another round – explicitly as possibilities, not rules.
This allows us to ask a much more precise question than „Does feedback work?“
What needs to be added to an experience for it to become a learning cycle?
Maybe experiencing and repeating a difficult conversation already changes something.
Maybe reflection is the crucial ingredient.
Maybe people need a concrete idea for what to try next.
Or perhaps explicit advice actually reduces exploration because people begin optimizing their behavior toward what they think the AI wants.
Those are empirical questions. And that is exactly why I like them.
The hardest question: does any of this transfer to humans?
There is an elephant in our virtual room: A person can become extraordinarily good at talking to DemocraGPT and still become furious with their uncle over dinner. So ultimately, our research cannot stop at whether participants perform differently inside the tool.
We need to distinguish between learning the bot and learning something about conflict.
That means asking whether the psychologically relevant mechanisms are sufficiently similar between AI and human interaction; whether the simulated dynamics feel consequential enough; whether people extract something abstract and portable from the experience; whether any changes persist; and whether effects can be distinguished from AI-specific effects.
For me, this is one of the intellectually most interesting parts of the project. Because it also pushes against some of the hype surrounding AI interventions. The relevant question is not simply whether people enjoy interacting with an AI or whether an AI can produce impressive personalized feedback. The relevant question is: What, precisely, can a human learn in interaction with a machine that still matters once the machine is gone?
Where we are now
As I write this in September 2026, DemocraGPT is very much a research project in development rather than a finished product. We have moved from the broad idea of „AI conversation training“ toward a much more specific architecture: a controlled arena in which generative AI functions as a challenging conversation partner and allows difficult interaction dynamics to be experienced, varied and studied.
We are developing the theoretical model behind those interactions. We are working on what different levels of difficulty should actually mean. We are investigating how to conceptualize and measure conversational quality. We are developing our learning and transfer model. And we are designing preliminary studies that will help us decide which levels, topics, experimental conditions and feedback mechanisms should eventually enter the larger evaluation.
This is exactly the stage at which research feels particularly alive to me: we have considerably better questions than when we started.
The bigger idea
I don’t think the future of democratic conversation lies in AI telling us how to talk to one another.
In fact, I find that idea rather unsettling.
But AI may offer something else.
It can create an interaction that responds to us. It can allow us to encounter resistance without involving another human being in our experiment. It can repeat a situation. It can systematically change aspects of it. And it can make dynamics visible afterwards that are extraordinarily difficult to notice while we are angry, defensive or busy formulating our next counterargument.
That makes generative AI an interesting experimental and experiential space.
The ambition behind DemocraGPT is therefore not to build a machine that knows how democracy should sound. It is to investigate whether a machine can create a sufficiently challenging space for humans to learn something about themselves, others, and what happens between them when disagreement gets difficult. And perhaps, eventually, to make it a little easier to stay in the conversation when it really counts.
DemocraGPT is a collaborative research project involving researchers from communication science, political science and computational social science. I will use this blog to share some of the questions, decisions and – inevitably – problems we encounter while developing and evaluating the project. Learn more about the project here.
