Language Models Are Capable of Metacognitive Monitoring and Control of Their Internal Activations
activation patternsadversarial attackai safetycognitive monitoringcognitive processesdimensionality reductionempirical evidencein-context learninginternal processesmetacognitionmetacognitive abilitiesneural activationneurofeedbacksemantic interpretability
Large language models (LLMs) can sometimes report the strategies they actually use to solve tasks, yet at other times seem unable to recognize those strategies that govern their behavior. This suggests a limited degree of metacognition