FineRS: Fine-grained Reasoning and Segmentation of Small Objects with Reinforcement Learning
attribute-level reasoningbounding boxcluttered contextscoarse-to-fine pipelineglobal semantic explorationhigh-resolution imagesinstruction-guided reasoninglocalized perceptual refinementlocate-informed retrospective rewardmulti-modal large language modelspixel-level segmentationreinforcement learning frameworksegmentation masksmall-scale targetsstate-of-the-art approachesvisual reasoning tasks
Multi-modal Large Language Models (MLLMs) have shown remarkable capabilities across a wide range of vision-language tasks. However, due to the restricted input resolutions, MLLMs face significant challenges in precisely understanding and localizing visual details in high-resolution images