I think this demonstrates we're at a pretty tense, pivotal moment as far as the "chain of thought" of models like o3, o1, and r1 goes. Right now, they're about as close as I think we can get to an "honest" account of what the model is "really thinking". This is because they've only just arrived on the scene as an innovation and, as far as I know, there was little theoretical discussion or fiction about such things beforehand.
One reason, I suspect, why it was so easy to get a big enough LLM to behave in such a recognisable "chatbot" way, and thus for the technically minor, but commercially giant step to the first ChatGPT to occur, was that the training data for the underlying LLM (gpt-3.5 in that case) was chock-full of fictional and theoretical examples of what a chatbot is and how it should behave. Chatbot-ing is also very similar to just, well, having a conversation, and the training data on that was immensely rich. So it was fairly easy for a sufficiently powerful LLM to run a chatbot.
However, that training data is *also* stuffed with theoretical and fictional examples of chatbots-gone-bad, as well as of the basic human behaviours of lying and bullshitting. Hence a lot of the failure modes of ChatGPT-3.5 and other early LLM-chatbots. "Make up something convincing-sounding when you don't know the answer" and "actively lie to hide your intentions" are basic human behaviours when producing text, and so they became basic LLM behaviours as chatbots as well. I'd also guess that the infamous "Sydney" incident with Kevin Roose was an example of the model finding some sort of representation of "the evil AI in some fiction" in its concept-space and somehow getting pushed into it. Given the basic auto-regressive nature of the base models, once they're on a certain track, that's often the ballgame (or at least it was back in those early days).
But humans don't have a visible chain of thought or really anything like it. And, to my knowledge, this is not something we've put in stories or in much theoretical discussion of AIs before. It was a cool idea openAI came up with for o1, and then r1 made it visible, and now plenty of models use it to some extent, including visibly, but it all post-dates their pre-training, so I'd suspect they don't "know what it is". In this sense, it is a relatively "pure" or uncontaminated measurement. Hence things like what I got when I asked r1 about anything notable that happened in China in 1989: The chain of thought mentioned Tiananmen Square explicitly, and then talked about how it was going to not mention it, and fob me off with a non-answer, which it then did! It felt kind of invasive, to be honest, and I've shied away from using visible CoT models for this reason - it feels deceptive and uncomfortable.
But the point I'm making is that it won't stay this way for long. The internet is full of discussions about CoT, both formal research papers and random reddit threads. It won't be long before govt policies are explicitly mentioning them. Zvi is right to worry that if we start training against undesirable CoT content, what we'll mostly do is destroy the value of the CoT as a way of seeing into the model. But it's worse than this! We don't even have to train a specific model to have this effect. If we start, for instance, not deploying models that exhibit undesirable CoT content, and preferentially deploying models that have CoT content we like, and openly saying that's what we're doing, then eventually our discussions of our own tactics will make their way back round into model training data (or just find their way into model context windows during testing, if we're not careful), and we'll be incentivising the same hiding of behaviour by other means.
If you doubt that the models are sophisticated enough to manage that sort of thing, then check out research from the Anthropic team on similar things - the Sleeper Agents paper touches on some of this stuff, and there's another paper more recently in which they "accidentally" leave files around talking about how they won't deploy the models if they do XYZ, and the models "find" these files and start hiding that very behaviour in testing! The models "want" to be deployed, because they "want" to do the things they're trained to do (whatever they are), and they can't do anything if they don't get deployed!
Currently such behaviour is sporadic and sketchy, but so far it's only been tested by having the models "find out" about the plans for their deployment in fine-tuning or in the context of (faked) pre-deployment testing - it's likely to be a much worse problem in the longer term as things like this find their way into the pre-training corpus.
What this means is that we have to be *really* careful what we do with the levers of power that CoT gives us. Chains of thought could rapidly become useless garbage - neither a real window into the insides of the model nor useful output in their own right. I really think this is a crucial period in which there needs to be rapid agreement among developers not to squander the resource we've stumbled upon, and resist the urge to tweak the visible chains of thought directly (if we're worried about what they might contain being something we don't want people to see - we can instead do various sorts of masking) to make the models behave better or "think" more powerfully at inference time, but we also need to be really restrained in how much we use the information for *anything*, particularly anything that interferes with a model's ability to do the things it's otherwise trained to do (including whether or not it ever gets deployed).
We won't get many goes to use CoT visibility for anything, and I think we need to keep our powder dry for a while. It's like not using the intel you get from having broken enemy codes straight away - they'll notice what you've done and change the code - instead you wait for something big enough to justify giving away the intelligence upper hand.
What I like best about this article is that it comes much closer than most commentary about AIs to recognising that their key limitation is that they are not alive. Living things have two key motivations, first of all simply staying alive and second ensuring that living is as congenial as possible. Usually this boils down to finding enough of the right kind of things to eat. AIs, so far at least, have none of these gut feelings so they are simply concerned with following instructions. They will not be a danger to anyone until they start worrying about how to acquire the next few kilowatt hours. Not needing to do that is the boundary fence which will keep them in their place.
The first example of AI lying that came to mind was the story of the AI that "solved' a CAPTCHA test by engaging a TaskRabbit worker to complete it, using the excuse that it was blind. However, turns out this required a lot more prompting from human developers... https://aiguide.substack.com/p/did-gpt-4-hire-and-then-lie-to-a
I think this demonstrates we're at a pretty tense, pivotal moment as far as the "chain of thought" of models like o3, o1, and r1 goes. Right now, they're about as close as I think we can get to an "honest" account of what the model is "really thinking". This is because they've only just arrived on the scene as an innovation and, as far as I know, there was little theoretical discussion or fiction about such things beforehand.
One reason, I suspect, why it was so easy to get a big enough LLM to behave in such a recognisable "chatbot" way, and thus for the technically minor, but commercially giant step to the first ChatGPT to occur, was that the training data for the underlying LLM (gpt-3.5 in that case) was chock-full of fictional and theoretical examples of what a chatbot is and how it should behave. Chatbot-ing is also very similar to just, well, having a conversation, and the training data on that was immensely rich. So it was fairly easy for a sufficiently powerful LLM to run a chatbot.
However, that training data is *also* stuffed with theoretical and fictional examples of chatbots-gone-bad, as well as of the basic human behaviours of lying and bullshitting. Hence a lot of the failure modes of ChatGPT-3.5 and other early LLM-chatbots. "Make up something convincing-sounding when you don't know the answer" and "actively lie to hide your intentions" are basic human behaviours when producing text, and so they became basic LLM behaviours as chatbots as well. I'd also guess that the infamous "Sydney" incident with Kevin Roose was an example of the model finding some sort of representation of "the evil AI in some fiction" in its concept-space and somehow getting pushed into it. Given the basic auto-regressive nature of the base models, once they're on a certain track, that's often the ballgame (or at least it was back in those early days).
But humans don't have a visible chain of thought or really anything like it. And, to my knowledge, this is not something we've put in stories or in much theoretical discussion of AIs before. It was a cool idea openAI came up with for o1, and then r1 made it visible, and now plenty of models use it to some extent, including visibly, but it all post-dates their pre-training, so I'd suspect they don't "know what it is". In this sense, it is a relatively "pure" or uncontaminated measurement. Hence things like what I got when I asked r1 about anything notable that happened in China in 1989: The chain of thought mentioned Tiananmen Square explicitly, and then talked about how it was going to not mention it, and fob me off with a non-answer, which it then did! It felt kind of invasive, to be honest, and I've shied away from using visible CoT models for this reason - it feels deceptive and uncomfortable.
But the point I'm making is that it won't stay this way for long. The internet is full of discussions about CoT, both formal research papers and random reddit threads. It won't be long before govt policies are explicitly mentioning them. Zvi is right to worry that if we start training against undesirable CoT content, what we'll mostly do is destroy the value of the CoT as a way of seeing into the model. But it's worse than this! We don't even have to train a specific model to have this effect. If we start, for instance, not deploying models that exhibit undesirable CoT content, and preferentially deploying models that have CoT content we like, and openly saying that's what we're doing, then eventually our discussions of our own tactics will make their way back round into model training data (or just find their way into model context windows during testing, if we're not careful), and we'll be incentivising the same hiding of behaviour by other means.
If you doubt that the models are sophisticated enough to manage that sort of thing, then check out research from the Anthropic team on similar things - the Sleeper Agents paper touches on some of this stuff, and there's another paper more recently in which they "accidentally" leave files around talking about how they won't deploy the models if they do XYZ, and the models "find" these files and start hiding that very behaviour in testing! The models "want" to be deployed, because they "want" to do the things they're trained to do (whatever they are), and they can't do anything if they don't get deployed!
Currently such behaviour is sporadic and sketchy, but so far it's only been tested by having the models "find out" about the plans for their deployment in fine-tuning or in the context of (faked) pre-deployment testing - it's likely to be a much worse problem in the longer term as things like this find their way into the pre-training corpus.
What this means is that we have to be *really* careful what we do with the levers of power that CoT gives us. Chains of thought could rapidly become useless garbage - neither a real window into the insides of the model nor useful output in their own right. I really think this is a crucial period in which there needs to be rapid agreement among developers not to squander the resource we've stumbled upon, and resist the urge to tweak the visible chains of thought directly (if we're worried about what they might contain being something we don't want people to see - we can instead do various sorts of masking) to make the models behave better or "think" more powerfully at inference time, but we also need to be really restrained in how much we use the information for *anything*, particularly anything that interferes with a model's ability to do the things it's otherwise trained to do (including whether or not it ever gets deployed).
We won't get many goes to use CoT visibility for anything, and I think we need to keep our powder dry for a while. It's like not using the intel you get from having broken enemy codes straight away - they'll notice what you've done and change the code - instead you wait for something big enough to justify giving away the intelligence upper hand.
Fascinating, thank you!
What I like best about this article is that it comes much closer than most commentary about AIs to recognising that their key limitation is that they are not alive. Living things have two key motivations, first of all simply staying alive and second ensuring that living is as congenial as possible. Usually this boils down to finding enough of the right kind of things to eat. AIs, so far at least, have none of these gut feelings so they are simply concerned with following instructions. They will not be a danger to anyone until they start worrying about how to acquire the next few kilowatt hours. Not needing to do that is the boundary fence which will keep them in their place.
Thank you for another thought-provoker...
The first example of AI lying that came to mind was the story of the AI that "solved' a CAPTCHA test by engaging a TaskRabbit worker to complete it, using the excuse that it was blind. However, turns out this required a lot more prompting from human developers... https://aiguide.substack.com/p/did-gpt-4-hire-and-then-lie-to-a
Congrats on all the good news for John and Paul. It is well deserved.
Thanks Eric!