Who Reasons in the Large Language Models?

Jie Shao (Nanjing University) · Jianxin Wu (Nanjing University)
circumstantial evidencediagnostic toolsempirical evidencefluent dialogueinternal behaviorsinterpretabilitymathematical reasoningmulti-head self-attentionoutput projection moduleoverfittingreasoning capabilitiesspecialized llmstargeted training strategiestransformer

Despite the impressive performance of large language models (LLMs), the process of endowing them with new capabilities