New Research on LLM "Thinking": You Can Now See Their Internal Workspace
A recent paper from Anthropic highlights an emergent internal 'workspace' within language models. These models generate silent words they can utilize for reporting, steering, and reasoning. A fascinating example is when a model is asked to evaluate '12 + 5 = 1'. It internally recognizes the incorrectness while still processing the problem, and the subsequent correction is essentially a narration of a decision already made.
This research has significant implications for the ongoing debate about whether LLMs truly reason or just autocomplete. It appears both perspectives hold some truth, and now we can observe this phenomenon directly. The tools developed allow us to see how much of a model's output, like grammar and common facts, bypasses this workspace, while complex, multi-step problems visibly utilize it.
I've put together a live viewer that integrates pre-fitted models, allowing you to watch this internal workspace in action, even before any output is generated. You can see the model processing information and planning responses. While this doesn't equate to consciousness, it's remarkable that this workspace isn't designed but rather emerges naturally in these models.
