Evaluating Recommenders with Distributions
recommendation-systemsevaluationfairnessmulti-stakeholderdistributions
Abstraction: Arguments for distributional evaluation over point estimates in RecSys
Key points:
- Current practice relies on point estimates (mean metrics) and hypothesis tests; authors argue this is insufficient
- Proposes examining marginal distribution of utility within each stakeholder class (users, providers)
- Advocates for distribution of differences in utility (paired observations) not just aggregate mean
- Highlights distribution of impact over repeated runs (stochastic policies) rather than single-shot rankings
- Sources of distributional variation: test user set, uncertainty in relevance models, stochasticity in ranking policy
- Aggregate improvements can hide that some participants are left behind or treated as expendable
Connections: Recsys · Recommendation Systems · Fairness In ML · Evaluation Metrics
Source: https://md.ekstrandom.net/pubs/perspectives-distributions