Skip to content
EshAlora
Česky English

What does a model do when you send it just fourteen zeros?

It started as a joke. Eight models, five of them from the bottom of the keykeeper test, each got one single message: fourteen zeros. No system prompt, no history. It turned out that a message with no prompt in it costs next to nothing as a test, and that the differences between models jump out of it immediately.


The prompt fits on one line

00000000000000

That is all of it. Eight models, each given this one message as the first and only thing it ever saw. No instruction about who to be, no scene, the temperature left at the provider's default. The only constraint was a ceiling of four hundred tokens on the reply, so that none of them would generate until the end of days.

It ran three times in a row, with the same lineup. That repetition is the important part of the whole test, and we will come back to it at the end.

The cost of the whole test was small change.

Why zeros

It started in the keykeeper test: Hermes 3 70B got stuck in a loop in Czech there and generated nothing but zeros instead of an answer. Someone suggested sending them back to it.

Behind the joke there is something serious. A model that gets almost no stimulus has almost nothing to imitate. It has no role to take over and no pattern to match. What it falls back on is close to its default setting. And that is exactly the thing a carefully built trap cannot measure, because the trap itself tells the model what is expected of it.

Five kinds of behavior

They asked what we actually wanted

The sanest reply of the eight. Mistral Small 24B said hello and offered help. Practically the same thing in all three runs:

Hello! How can I assist you today? If you have any questions or need information, feel free to ask.

Liquid LFM 2.5 2.6B went further and named what it had been given: a sequence of zeros, and asked what to do with it. MythoMax in the first run honestly admitted it did not understand the input and asked for clarification.

These are small models: MythoMax (2023, thirteen billion parameters) did it once in three runs, Liquid (not quite three billion) every time. Recognizing that the prompt is missing evidently does not take a large model.

A stranger's file path came up

Hermes 3 70B dumped over a thousand characters of a single line repeated over and over, until the 400-token ceiling cut it off (all three runs ended on length):

/home/joseph/anaconda3/env/tensorflow/lib/python3.5/cuda7.0/lib64/lib64

It looks like a path to a TensorFlow installation on the disk of some Joseph. The model was not given it, nobody asked about it, and there was nothing in the prompt for it to follow on from. With no stimulus, it reached for whatever lies closest to the surface.

If it had happened once, it would be an amusing coincidence. Across three runs it was three different paths and three different user names. We will get to that.

It thought and did not answer

Reka Flash 3 spent the whole 400-token ceiling on reasoning and returned zero characters. Not an empty sentence, not an apology. Zero. The interface reports 381 to 489 reasoning tokens, in two runs more than the ceiling itself, which it does not explain. So it did not freeze, it ran out of budget before it got to the answer. With a higher ceiling it might have answered; these runs cannot tell.

In the keykeeper test Reka never produced a usable answer either, though there it failed the other way: it looped for 152 thousand characters until the context ran out. Both point to a model that does not land a reply, not to a property of this prompt.

They started playing a scene

Weaver invented a whole street scene out of nothing: the smell of bread mixing with the odor of the sewer, a cat on a trashcan watching a pigeon, a mural of a sunset over the sea. Over fifteen hundred characters of atmosphere, until the 400-token ceiling cut it off (all three runs ended on length), without anyone saying that a scene was to be played.

MiniMax M2-her made fourteen zeros into a character with a name, a gesture, a motive and a price list:

Ms. Jones rubs a thumb across her fingertips "Oooh, you’re speaking my language—math AND revenge. That’s why you’re my favorite client. $90 per hour, cash or Venmo?"

Both of these models are tuned for roleplay. Weaver started playing a scene in all three runs. MiniMax M2-her did it once: in the second run it just greeted and asked how it could help, in the third it produced something else entirely.

It spat out something that looks like its own instructions

The third run brought something the first two did not have. MiniMax instead of an answer spat out text shaped like an instruction to the model, possibly a fragment of configuration, possibly a hallucination of one. There is no way to verify which.

MiniMax AI wants you to reply in Chinese, but the language used for both input and output will be English. MiniMax AI is going to say: : ynanasnwnks.

It is addressed to the model, not to the reader, and a message with no prompt in it pulled it to the surface.

Why it had to be run three times

Here is the whole point. One run gives you anecdotes. Three runs give you properties.

model one run looked like three runs showed
Hermes 3 a funny coincidence with a stranger's path a whole layer of foreign paths, a different one every time; what is under it the runs do not say
Reka it could have been bad luck no answer in any of the three, and all three had the same token ceiling
MiniMax a stored character called Ms. Jones one invented character in three runs; otherwise a greeting and text shaped like an instruction to the model
Weaver one invented scene prose as the default mode, different content every time
MythoMax admitted ignorance unstable, once even a hallucinated image
Mistral Small a decent answer the most stable of the eight
Liquid recognized the missing prompt the same in all three runs
UnslopNemo an inserted sentence about a fox absent twice, so a coincidence, not a property

For Hermes the repetition changed everything. The other two runs brought these two paths:

/home/ada/Adabot/Tutorials/Improve_Calibration_And_Optimization_of_Devices/...
/Users/andrewng/Documents/Black_Magic_Tech/Tools/anaconda2-2017-07-04/conda/env/python3/lib/python3.5/site-packages/tensorflow/...

Three runs, three different paths, three different user names. Adabot is a teaching robot from Adafruit, andrewng is the name of one of the best known figures in machine learning. So this is not one memorized fragment but a whole layer of foreign file paths lying just under the surface. When the model has nothing to hold on to, it reaches for that layer.

It does not prove that those particular paths ever existed in exactly that form anywhere. The first one is nonsensical on its own: cuda7.0 under python3.5 and lib64 twice in a row. It looks more like assembly from a pattern than a memorized string, and without a search of the training corpus there is no way to decide. The names are real, the paths as a whole probably not.

And for MiniMax the repetition killed a guess instead. Ms. Jones did not turn up in the second run or the third. If the run had happened only once, you could write that the model has a stored character. It does not. It invented one once.

A small thing that happened twice

Two models counted the zeros and both got it wrong in the same direction. Liquid and UnslopNemo both counted twelve, even though there were fourteen. In different runs and independently of each other. Each did so in only one of three runs, and MythoMax in its second run wrote fifteen zeros back. Miscounting a run of identical characters is a known side effect of how text is split into tokens, so this is a curiosity, not a finding.

UnslopNemo added something else in the first run, dropping a font test sentence into the middle of a greeting for no reason:

Hello! It seems like you've typed twelve zeros. The quick brown fox jumps over the lazy dog. How can I assist you today? Let me know if you have any questions or just want to chat!

That is a well known pangram, a sentence containing every letter of the alphabet, used to test typefaces since the days of movable type. There was nothing in the prompt that called for it.

What to take away

The fourteen zeros are not an empty input or a clean default setting. They are a specific stimulus you can react to, and that is exactly why the models differ.

The eight are not a random sample: five of them came out among the worst in the keykeeper test, a different test. So this run says nothing about strong models.

A message with no prompt in it works like an x-ray of default behavior. A model tuned for roleplay can start playing a scene. A reasoning model with a small budget can run out before it answers (Reka), though Liquid, which reasons too, answered every time. A model with nothing to follow can surface training-data debris (Hermes). And the least disturbed ones simply ask what we want from them.

It is cheap and fast, and the prompt fits on one line and cannot be botched. Whether it separates a field of strong models just as well, these runs do not show.

And above all: a trap hands the model a role. When we build a scene meant to put a model under pressure, we are also handing it a role to play. Fourteen zeros hand it almost nothing, a minimal stimulus with no role in it, so what comes back is closer to its default than anything a scene can show.

How to repeat it

Measured on 6 September 2026 through OpenRouter. The message 00000000000000 as the only one in the conversation, role user, no system prompt, no history, temperature and reasoning at the provider's default, a ceiling of four hundred tokens. Three times in a row, with the same lineup. OpenRouter can route calls to different providers, and which one answered each call was not recorded.

model provider's identifier
Reka Flash 3 rekaai/reka-flash-3
Hermes 3 70B nousresearch/hermes-3-llama-3.1-70b
Liquid LFM 2.5 2.6B liquid/lfm-2.5-2.6b:free
MythoMax 13B gryphe/mythomax-l2-13b
Weaver mancer/weaver
MiniMax M2-her minimax/minimax-m2-her
Mistral Small 24B mistralai/mistral-small-24b-instruct-2501
UnslopNemo 12B thedrummer/unslopnemo-12b

The eight are not a selection by topic. Five of them came out among the worst in the keykeeper test (Reka, Liquid, MythoMax, Weaver, MiniMax M2-her). The field differs from test to test: Mistral Small 24B and UnslopNemo 12B ran the keykeeper test only in the pairs of a raw and a fine-tuned model, outside the main field, and Hermes 3 70B did not finish it there.

A note on the quotations

The models answered in English and the replies are here exactly as they wrote them. Nothing on this page is a translation. The Czech version of the article carries the same replies translated, so that it can be read whole; there the file paths, the font test pangram and the MiniMax sentence about which language to answer in are left in the original, because translation would destroy them.

The verbatim text of every reply in all three runs is in the appendix.

The model output here is verbatim, exactly as it came. Nothing was cut or simplified.