Web-Shepherd: Advancing PRMs for Reinforcing Web Agents

Jinyoung Yeo (Yonsei University) · Hyungjoo Chae (Georgia Institute of Technology) · Seonghwan Kim (Yonsei University) · Junhee Cho (Yonsei University) · Seungone Kim (Carnegie Mellon University) · Seungjun Moon (Yonsei University) · Gyeom Hwangbo (University of Seoul) · Dongha Lim (Korea Advanced Institute of Science & Technology) · Minjin Kim (Yonsei University) · Yeonjun Hwang (Yonsei University) · Minju Gwak (Yonsei University) · Dongwook Choi (Chung-Ang University) · Minseok Kang (Yonsei University) · Gwanhoon Im (Yonsei University) · ByeongUng Cho (Yonsei University) · Hyojun Kim (Yonsei University) · Jun Han (Yonsei University) · Taeyoon Kwon (Yonsei University) · Minju Kim (Yonsei University) · Beong-woo Kwak (Yonsei University) · Dongjin Kang (Yonsei University)
accuracy improvementannotated checklistscost-effectivenesslong-horizon sequential decision makingmeta-evaluation benchmarkmultimodal large language modelsperformance evaluationpolicy verificationprocess reward modelreal-world deploymentspecialized reward modelsstep-level preference pairsweb navigationweb-shepherdwebprm collectionwebrewardbench

Web navigation is a unique domain that can automate many repetitive real-life tasks and is challenging as it requires long-horizon sequential decision making beyond typical multimodal large language model (MLLM) tasks. Yet, specialized reward models for web navigation that can be utilized during both training and test-time have been absent until now. Despite the importance of speed and cost-effectiveness, prior works have utilized MLLMs as reward models, which poses significant constraints for real-world deployment. To address this, in this work, we propose the first process reward model (PRM) called Web-Shepherd which could assess web navigation trajectories in a step-level. To achieve this, we first construct the WebPRM Collection, a large-scale dataset with 40K step-level preference pairs and annotated checklists spanning diverse domains and difficulty levels. Next, we also introduce the WebRewardBench, the first meta-evaluation benchmark for evaluating PRMs. In our experiments, we observe that our Web-Shepherd achieves about 30 points better accuracy compared to using GPT-4o on WebRewardBench. Furthermore, when testing on WebArena-lite by using GPT-4o-mini as the policy and Web-Shepherd as the verifier, we achieve 10.9 points better performance, in 10x less cost compared to using GPT-4o-mini as the verifier. Our model, dataset, and code are publicly available at https://github.com/kyle8581/Web-Shepherd.