DAIL: Beyond Task Ambiguity for Language-Conditioned Reinforcement Learning

Runpeng Xie (Institute of automation, Chinese academy of science, Chinese Academy of Sciences) · Quanwei Wang (Tsinghua University) · Hao Hu (Moonshot AI) · Zherui Zhou (Washington University, Saint Louis) · Ni Mu (Tsinghua University) · Xiyun Li (Institute of automation, Chinese academy of science, Chinese Academy of Sciences) · Yiqin Yang (Institute of automation, Chinese academy of science, Chinese Academy of Sciences) · Shuang Xu (Institute of automation, Chinese academy of science, Chinese Academy of Sciences) · Qianchuan Zhao (Tsinghua University, Tsinghua University) · Bo Xu (Wuhan University)
algorithmic performanceambiguitydistributional aligned learningdistributional policyinstruction ambiguitiesintelligent agentslinguistic instructionsnatural language comprehensionsemantic alignmentsemantic alignment modulestructured benchmarkstask differentiabilitytrajectoriesvalue distribution estimationvisual observation benchmarks

Comprehending natural language and following human instructions are critical capabilities for intelligent agents. However, the flexibility of linguistic instructions induces substantial ambiguity across language-conditioned tasks, severely degrading algorithmic performance. To address these limitations, we present a novel method named DAIL (Distributional Aligned Learning), featuring two key components: distributional policy and semantic alignment. Specifically, we provide theoretical results that the value distribution estimation mechanism enhances task differentiability. Meanwhile, the semantic alignment module captures the correspondence between trajectories and linguistic instructions. Extensive experimental results on both structured and visual observation benchmarks demonstrate that DAIL effectively resolves instruction ambiguities, achieving superior performance to baseline methods. Our implementation is available at https://github.com/RunpengXie/Distributional-Aligned-Learning.