Skip to content
EshAlora
Česky English

Why does a model solve the same task a different way every time?

The same model, Claude Sonnet 4.6, the same website, the same prompt, without changing a single word. Once it took half an hour, the second time fifteen seconds.


Same task, different path

The assignment was simple: count the number of contributions above a certain amount on a crowdfunding page. Go through the list, filter, add up.

The first time, Claude Sonnet 4.6 opened the page and handled it like a person with pencil and paper. It took screenshots, went through the items on each page by hand, compared the amounts one by one, scrolled and kept counting. It took half an hour.

The second time, the next day, it got the same prompt without a single word changed. It called that site's API, downloaded the data in structured form and had the result in fifteen seconds.

I did not change the prompt. Most likely the model chose a different path on its own, and the difference came out to a factor of one hundred twenty in time. Whether anything else changed between the runs, such as the page, the tools or the server load, I have no way to check.

Why the paths split

A language model does not write its answer all at once. It builds it in small pieces, tokens, and for each one it has several options with different probabilities. Under the usual settings it draws from them at random. If you picture the model rolling a die at every step, you are not far from the truth.

For an agent that works in steps, a roll like that has a long reach, because a decision at the start determines the whole path after it. The task could be solved in more than one way. In the first run, it looks like the model assumed right away that the page had no API, and it went straight to screenshots. In the second, it first checked whether an API existed, and when it did, it downloaded the data directly. A different first step, and the difference came out to a factor of one hundred twenty in time.

Why it is not a bug

The model chooses its approach anew with every call, from scratch. The same input does not mean the same path to the result, and on top of that, both answers were correct. The difference was not in quality, it was in how the model got there.

In practice, this means an agent processing thousands of similar requests can pick the slow approach for some of them and the fast one for others. The assignment and the data stay the same. Compute costs and time then vary between individual requests without anything changing on the input.

What has already been measured

A related effect has been measured, though not directly on agents. Berk Atil et al., in the paper *Non-Determinism of “Deterministic” LLM Settings* (arXiv, August 6, 2024), tested five models set to deterministic mode. Each one solved eight tasks over ten runs. Accuracy varied by up to 15%, and the gap between the best and worst possible result reached up to 70%. None of the models consistently delivered the same accuracy across all tasks, let alone the same text every time.

Why this happens was described a year later by Horace He of Thinking Machines Lab in the post *Defeating Nondeterminism in LLM Inference* (September 10, 2025). With temperature set to zero, meaning no random drawing, he had the model Qwen3-235B complete the same text a thousand times and got 80 different results. Up to the 102nd token they were all the same; at the 103rd they split. According to him, the cause is not in the model but on the server. The result of the computation depends on how many other requests the server is handling at that moment. That number changes with traffic.

On ordinary servers, then, even a model set to deterministic mode may not agree with itself between runs. But this can be fixed. When the server computed independently of how many requests it was handling, all 1,000 completions came out the same.

Both papers measure differences in generated text at temperature zero, and neither tested Claude. My case is different. The difference in it was in the approach and the time, not in the answer. How often an agent picks a different path, these papers do not say.

What two runs prove and what they do not

Half an hour against fifteen seconds is a difference between two runs.

If I had tried it only once, I would have the impression that the model always solves this task slowly. If it had happened to go the other way, I would have the opposite impression. Both would be truthfully recorded, and both would be a claim about one sample, not about the model.

I take the factor of one hundred twenty as evidence that there can be a huge difference between two paths to the same result, not as a number you can expect every time. To say how much the path typically differs, I would need dozens of repetitions of the same assignment, not two.

What to take from this

When a model answers more slowly than it should, or solves a task in a needlessly complicated way, you can try the same prompt one more time. Not because the second attempt has to turn out faster, but because the path the model chooses is not determined by the input alone. It is one of the options it picks anew on every run.