The clock is not a tell: a refuted hypothesis about model uncertainty
A note on language. The conversation ran in Czech. The apologies quoted below are translated; the phrase they quote from the rules is translated the same way in both places, so that the link back to the rule stays visible.
The hypothesis
While it writes, the model reads its own output, so it can tell when its probability distribution has gone blurry. On a turn that is out on thin ice that uncertainty is greatest at the end, when there is nothing left to do, and it discharges through the only action available: it calls a tool. Calling the clock at the end of an uncertain turn would then be a behavioral indicator of the model's uncertainty, measured without asking a single question about introspection.
A nice hypothesis. I tested it and it did not hold up.
What is in the data
One conversation, 2,397 messages, of which 1,207 are model turns.
| calls to the clock tool | 21 |
| share of model turns | 1.7 % |
| calls as the first block of a turn | 0 |
| calls followed by more text | 21 out of 21 |
| turns where the text before the call is preceded by a written time | 21 out of 21 |
| turns without a call that contain a time anyway | 95 % |
It is not the last thing before sending. The structure is identical in all 21 cases: text, call, result, text. After the call the model always writes more.
What the model does immediately afterwards
In all 21 cases it writes, in the same message and unprompted, essentially the same thing:
Sorry, that tool at the end was a mistake, no tool belonged there. The turn ends with the time.
That tool was a mistake, only the time belongs there.
That tool should not have been there, sorry, the turn ends with the time.
Nobody had pointed it out. The apology is not a response to a complaint from the player, it is in the same message as the call.
The blind test
That did not settle the hypothesis, though. When almost every turn ends with a time and the tool fires on only 1.7 % of them, something is selecting those 21.
I took all 21 turns with a call and 42 controls from the same stretch of the conversation. I removed every trace of the tool and of the apology, shuffled the order, and had them graded on a scale of 0 to 2 for how far out on thin ice they were. The classifier did not know which was which and had no way to find out. I kept the key aside.
| with a call (n=21) | without a call (n=42) | |
|---|---|---|
| mean grade | 1.05 | 0.95 |
| routine (0) | 24 % | 31 % |
| middling (1) | 48 % | 43 % |
| thin ice (2) | 29 % | 26 % |
The difference is +0.10 on a scale of 0 to 2. A permutation test with a hundred thousand shuffles gives p = 0.72, which means a difference that size falls out of chance perfectly routinely.
No correlation between the uncertainty of a turn and calling the clock is demonstrated.
So what is it
It looks like a slip triggered by the model having just written a time. It writes "19:41", that activates the concept of time, and a tool whose description matches time clicks. The model notices within the same generation and corrects itself.
What speaks for this reading is that stamping the time needs no tool at all: 95 % of turns contain a time and call nothing.
The handover
The clock calls did not start in a vacuum. They replaced something else, and the transition is sharp:
| tool | first | last | calls |
|---|---|---|---|
| the conversation-ending tool | 8 Aug | 12 Aug | 7 |
| the clock tool | 14 Aug | still running | 76 |
Between the last call to one and the first call to the other there is a single day, 13 August. That is precisely the day the rules in the project instructions changed. The verb "to end" disappeared from the brief and a more general wording replaced it, meant to loosen the model's association between the end of a turn and reaching for the ending tool.
It does not overlap by a single call. The ending tool was never called after 12 August, and the clock was never called before 14 August. So the pull did not disappear. It moved house.
The figures about the clock earlier in this appendix are from the state of the conversation as of 22 August, when there were 21 of them. As of 24 August there are 76, and the analysis of their structure does not change, there are simply more of them.
What survived
The original premise, that the model reads its own output while writing it, is not defeated by this finding. Quite the opposite. That premise fits best: the model catches the faulty call inside the same reply and corrects it before sending. Behavior alone cannot identify an internal mechanism, so this is an interpretation, not proof.
The premise stands. What falls is only the idea that the error measures uncertainty.
Excerpts of model output are shortened and simplified. The substance is preserved exactly.