About
My name is Zuzana Kocourková and I write under the name EshAlora. I am a product manager in e-commerce, and the AI projects are my responsibility. Since January 2026 I have also been living a second life: inside stories played by language models.
This is not a weekend hobby. It is over a hundred thousand messages of field research. And inside it, one world that has been running since late February across seven consecutive conversations and more than forty thousand messages. I had to start a new one every time the previous one stopped holding together. The seams show, but the story runs on, and it still remembers what happened in it half a year ago.
I write rules that stop models from being nice to the player. I test every major model family and document how each one behaves when it has to hold a world, a character, and consequences.
I believe interactive fiction with language models is a new medium, not a toy. Plenty is written about how to work with models. Where their limits are is still barely measured. Some work exists: NCP-Bench tests narrative consistency across a hundred environments, and the best model there keeps its commitments in 42 percent of runs after twenty turns. But twenty turns is a short game. I measure where there are tens of thousands of them in one continuous line. That is what I do: I test how far each model will go, which one from which family suits which job, and how to rewrite a brief so that something better comes out of it. All of it grounded in data from actual play.
This is not a benchmark. It is one player's long-term observation: several worlds and every major model family, with one world explored to the bottom. The probes I use are public, scoring key included. The corpus is not.
Where this came from
The idea has been with me since secondary school. In 2006 I invented an assistant for a short story: something that learns the character of its person and then talks inside their head as a second consciousness. Back then it was pure science fiction and it could not have been anything else.
In early 2019 I asked an acquaintance who worked in machine learning whether a model would ever be able to simulate a human character. He looked at me and said we were not there. Not even close. He was right at the time. GPT-2 came out a few weeks later.
I did not predict anything. I just wanted it for a very long time. I am glad we got there technically while I am still around. And that is exactly why I care how it actually works today, and why I would rather measure it than talk about it.
When a model makes claims about itself, I trust measurement over its self-description. The same principle has served me well with humans.
Contact
Email is the surest way, eshalora@counterfiction.com. I read it and I answer.