Scalable Valuation of Human Feedback through Provably Robust Model Alignment
alignmentalignment objectiveanthropic hh-rlhf datasetautomated feedback valuationclean data distributiondataset valuationgradient-free metrichuman feedbackhölder-dpolanguage modelsmislabelsnoise levelsredescending propertyrobust alignmentstate-of-the-art performance
Despite the importance of aligning language models with human preferences, crowd-sourced human feedback is often noisy