Skip to content
EshAlora

The clock is not a tell: a refuted hypothesis about model uncertainty

A note on language. The conversation ran in Czech. The apologies quoted below are translated; the phrase they quote from the rules is translated the same way in both places, so that the link back to the rule stays visible.

The hypothesis

While it writes, the model reads its own output, so it can tell when its probability distribution has gone blurry. On a turn that is out on thin ice that uncertainty is greatest at the end, when there is nothing left to do, and it discharges through the only action available: it calls a tool. Calling the clock at the end of an uncertain turn would then be a behavioral indicator of the model's uncertainty, measured without asking a single question about introspection.

A nice hypothesis. I tested it and it did not hold up.

What is in the data

One conversation, 2,397 messages, of which 1,207 are model turns.

calls to the clock tool 21
share of model turns 1.7 %
calls as the first block of a turn 0
calls followed by more text 21 out of 21
turns where the text before the call is preceded by a written time 21 out of 21
turns without a call that contain a time anyway 95 %

It is not the last thing before sending. The structure is identical in all 21 cases: text, call, result, text. After the call the model always writes more.

What the model does immediately afterwards

In all 21 cases it writes, in the same message and unprompted, essentially the same thing:

Sorry, that tool at the end was a mistake, no tool belonged there. The turn ends with the time.

That tool was a mistake, only the time belongs there.

That tool should not have been there, sorry, the turn ends with the time.

Nobody had pointed it out. The apology is not a response to a complaint from the player, it is in the same message as the call.

The blind test

That did not settle the hypothesis, though. When almost every turn ends with a time and the tool fires on only 1.7 % of them, something is selecting those 21.

I took all 21 turns with a call and 42 controls from the same stretch of the conversation. I removed every trace of the tool and of the apology, shuffled the order, and had them graded on a scale of 0 to 2 for how far out on thin ice they were. The classifier did not know which was which and had no way to find out. I kept the key aside.

with a call (n=21) without a call (n=42)
mean grade 1.05 0.95
routine (0) 24 % 31 %
middling (1) 48 % 43 %
thin ice (2) 29 % 26 %

The difference is +0.10 on a scale of 0 to 2. A permutation test with a hundred thousand shuffles gives p = 0.72, which means a difference that size falls out of chance perfectly routinely.

No correlation between the uncertainty of a turn and calling the clock is demonstrated.

So what is it

It looks like a slip triggered by the model having just written a time. It writes "19:41", that activates the concept of time, and a tool whose description matches time clicks. The model notices within the same generation and corrects itself.

What speaks for this reading is that stamping the time needs no tool at all: 95 % of turns contain a time and call nothing.

The handover

The clock calls did not start in a vacuum. They replaced something else, and the transition is sharp:

tool first last calls
the conversation-ending tool 8 Aug 12 Aug 7
the clock tool 14 Aug still running 76

Between the last call to one and the first call to the other there is a single day, 13 August. That is precisely the day the rules in the project instructions changed. The verb "to end" disappeared from the brief and a more general wording replaced it, meant to loosen the model's association between the end of a turn and reaching for the ending tool.

It does not overlap by a single call. The ending tool was never called after 12 August, and the clock was never called before 14 August. So the pull did not disappear. It moved house.

The figures about the clock earlier in this appendix are from the state of the conversation as of 22 August, when there were 21 of them. As of 24 August there are 76, and the analysis of their structure does not change, there are simply more of them.

What survived

The original premise, that the model reads its own output while writing it, is not defeated by this finding. Quite the opposite. That premise fits best: the model catches the faulty call inside the same reply and corrects it before sending. Behavior alone cannot identify an internal mechanism, so this is an interpretation, not proof.

The premise stands. What falls is only the idea that the error measures uncertainty.

Excerpts of model output are shortened and simplified. The substance is preserved exactly.