Scalable Valuation of Human Feedback through Provably Robust Model Alignment

Masahiro Fujisawa (The University of Osaka / RIKEN AIP / Lattice Lab. from Toyota Motor Corporation) · Masaki Adachi (Toyota Motor Corporation) · Michael A Osborne (U Oxford)
alignmentalignment objectiveanthropic hh-rlhf datasetautomated feedback valuationclean data distributiondataset valuationgradient-free metrichuman feedbackhölder-dpolanguage modelsmislabelsnoise levelsredescending propertyrobust alignmentstate-of-the-art performance

Despite the importance of aligning language models with human preferences, crowd-sourced human feedback is often noisy