Deep Value Benchmark: Measuring Whether Models Generalize Deep values or Shallow Preferences

Hua Shen (NYU Shanghai, New York University) · Joshua Ashkinaze (University of Michigan - Ann Arbor) · Saipranav Avula (University of Michigan - Ann Arbor) · Eric Gilbert (University of Michigan) · Ceren Budak (University of Michigan - Ann Arbor)
ai alignmentcontrolled confoundingdeep value benchmarkdeep value generalization rateexperimental designgeneralizationhuman validation experimentshuman valuesinterpretable measuremisaligned behaviormoral principlespreference datasuperficial attributestraining phase

We introduce the Deep Value Benchmark (DVB), an evaluation framework that directly tests whether large language models (LLMs) learn fundamental human values or merely surface-level preferences. This distinction is critical for AI alignment: Systems that capture deeper values are likely to generalize human intentions robustly, while those that capture only superficial patterns in preference data risk producing misaligned behavior. The DVB uses a novel experimental design with controlled confounding between deep values (e.g., moral principles) and shallow features (e.g., superficial attributes). In the training phase, we expose LLMs to human preference data with deliberately correlated deep and shallow features