FineRS: Fine-grained Reasoning and Segmentation of Small Objects with Reinforcement Learning

Huchuan Lu (Dalian University of Technology) · Lu Zhang (Dalian University of Technology) · Jiazuo Yu (Dalian University of Technology) · Haomiao Xiong (Dalian University of Technology) · Ping Hu (University of Electronic Science and Technology of China) · Yunzhi Zhuge (Dalian University of Technology) · You He (Tsinghua University, Tsinghua University)
attribute-level reasoningbounding boxcluttered contextscoarse-to-fine pipelineglobal semantic explorationhigh-resolution imagesinstruction-guided reasoninglocalized perceptual refinementlocate-informed retrospective rewardmulti-modal large language modelspixel-level segmentationreinforcement learning frameworksegmentation masksmall-scale targetsstate-of-the-art approachesvisual reasoning tasks

Multi-modal Large Language Models (MLLMs) have shown remarkable capabilities across a wide range of vision-language tasks. However, due to the restricted input resolutions, MLLMs face significant challenges in precisely understanding and localizing visual details in high-resolution images