Taming Hyperparameter Sensitivity in Data Attribution: Practical Selection Without Costly Retraining

Han Zhao (University of Illinois, Urbana Champaign) · Yuzheng Hu (University of Illinois Urbana-Champaign) · Jiaqi Ma (University of Illinois Urbana-Champaign) · Weiyi Wang (University of Michigan Ann Arbor) · Junwei Deng (University of Illinois Urbana-Champaign) · Shiyuan Zhang · Xirui Jiang (University of Michigan - Ann Arbor) · Runting Zhang (University of Michigan - Ann Arbor)
computationally-cheap validation metricsdata attributiondata-centric applicationsempirical studyhyperparameter sensitivityhyperparameter tuninginfluence function methodslightweight proceduremethod developmentmodel retrainingpractical applicationregularization termstandard data attribution benchmarkstheoretical analysistuning strategies

Data attribution methods, which quantify the influence of individual training data points on a machine learning model, have gained increasing popularity in data-centric applications in modern AI. Despite a recent surge of new methods developed in this space, the impact of hyperparameter tuning in these methods remains under-explored. In this work, we present the first large-scale empirical study to understand the hyperparameter sensitivity of common data attribution methods. Our results show that most methods are indeed sensitive to certain key hyperparameters. However, unlike typical machine learning algorithms