On Optimal Steering to Achieve Exact Fairness
affine steeringbias inbias outcost-sensitive riskdemographic parityfair machine learningfeature distributionsgroup-fair outcomesideal distributioninternal representationskl-divergencemulti-class classificationoptimization programparametric familiesutility improvement
To fix the `bias in, bias out' problem in fair machine learning, it is important to steer feature distributions of data or internal representations of Large Language Models (LLMs) to \emph{ideal} ones that guarantee group-fair outcomes. Previous work on fair generative models and representation steering could greatly benefit from provable fairness guarantees on the model output. We define a distribution as \emph{ideal} if the minimizer of any cost-sensitive risk on it is guaranteed to have exact group-fair outcomes (e.g., demographic parity, equal opportunity)