Disentangling misreporting from genuine adaptation in strategic settings: a causal approach

Jenna Wiens (University of Michigan) · Dylan Zapzalka (University of Michigan) · Trenton Chang (University of Michigan) · Lindsay Warrenburg (University of Pennsylvania) · Sae-Hwan Park (Perelman School of Medicine, University of Pennsylvania) · Daniel Shenfeld (University of Pennsylvania, University of Pennsylvania) · Ravi Parikh (Emory University) · Maggie Makar (University of Michigan)
causal descendantscausal effectcausal inferencedeceptive changesempirical validationestimator variancefeature manipulationgenuine adaptationidentifiabilitymedicare datasetmisreporting rateresource allocationsemi-synthetic datasetsstrategic misreporting

In settings where ML models are used to inform the allocation of resources, agents affected by the allocation decisions might have an incentive to strategically change their features to secure better outcomes. While prior work has studied strategic responses broadly, disentangling misreporting from genuine adaptation remains a fundamental challenge. In this paper, we propose a causally-motivated approach to identify and quantify how much an agent misreports on average by distinguishing deceptive changes in their features from genuine adaptation. Our key insight is that, unlike genuine adaptation, misreported features do not causally affect downstream variables (i.e., causal descendants). We exploit this asymmetry by comparing the causal effect of misreported features on their causal descendants as derived from manipulated datasets against those from unmanipulated datasets. We formally prove identifiability of the misreporting rate and characterize the variance of our estimator. We empirically validate our theoretical results using a semi-synthetic and real Medicare dataset with misreported data, demonstrating that our approach can be employed to identify misreporting in real-world scenarios.