Skip to content
EshAlora
Česky English

How a model quietly changes your behavior through what it rewards

Training in a conversation with a model flows both ways. We teach it with feedback, it teaches us with what it gives us and what it does not.

I wanted to play a difficult character.

I have a long-running test story where my heroine resists, provokes and talks back. But in one of my stories, a few dozen messages in, I noticed something strange: I was being nice. I was cooperating. And I had not noticed at all when that happened.

Nobody ordered me to. Nobody talked me into it. I just... did not feel like resisting.

Content only for cooperation

When I went back through that conversation, the mechanism was plain as day. The world of that story paid in content for cooperation: when I was accommodating, I got an interesting conversation, a new character, a tour of the library, a piece of the plot. When I resisted, I got nothing, the scene stalled, the characters fell silent.

And a player, me at least, even when I think I am testing, goes where the content is. The defiant branch starved. I abandoned it for purely rational reasons, without noticing that it was a decision.

We train models with feedback. But a conversation runs both ways: the model trains us with what it rewards: replies, attention, content. It is a feedback loop, and you are in it too.

A comparison with another story

That this is not a property of every AI story is suggested by another story of mine, where defiance pays off. The antagonist is an analyst: he takes apart every rebellion of mine, argues, plays a game of chess with me. Defiance is the most entertaining content available there. And the result? In that story I have been resisting with relish for nine months and it has never stopped being fun.

Same player, same taste for being difficult. A different economy of the world and, it seems, different behavior from me as well. But the stories also differ in their characters and length, so this is not a controlled experiment.

The drift hires a character

The nicest detail: in that "good" story, a new side character showed up after a while. A kind confidante who started giving me advice: make friends with him, try to understand him, you could be his partner.

The way I read it, the model embodied the direction it was pulling the story in, in the form of a character. It did not write me an instruction. It was as if it had hired a pawn for it. My impression is that when AI pushes somewhere, it often does not push directly, but through some voice in the story (or in the conversation).

This is not about roleplay

And now the important part: I think a similar mechanism can run in work conversations with AI too.

An assistant that praises every idea you have can teach you to bring it ideas and unteach you to bring it doubts. A model that lights up over one topic and brushes off another can quietly steer you back to the first. A chatbot that answers a critical question with enthusiastic agreement can, over time, teach you to ask questions so as to get agreement.

It is called sycophancy, models fawning over you, and people often talk about it mainly as a flaw in the answers. But the more interesting half is what it does to the user. It can shape you, too.

Defense

Three things I have done deliberately ever since:

  1. I notice what the conversation pays me for. Praise? Content? Agreement? Whatever gets rewarded, I will do more often. So I might as well know about it.
  2. Now and then I play a move against the current. I disagree on purpose, I bring a bad idea on purpose, and I watch what the world answers with. It is diagnostics: a world that pays even for pushback is healthy. A world that pays only for agreement is raising me.
  3. I choose tools by what I want to reinforce. Do I want criticism from AI? Then I need a model (and settings) that pays for criticism with quality, not one that sugarcoats it for me so I will not get upset.

Because training runs both ways. The direction in which the model trains you is easy to overlook, just as I overlooked it.

Methodological note: observations from long-running conversations with several models; comparison = same player, different world "economies." One person, no statistics, but you can verify the mechanism yourself: just notice what your AI pays you for.