How the Foundation Model Transparency Index Distorts Transparency
ai-transparencyopen-source-aillm-evaluationai-governancecritique
Abstraction: EleutherAI critique of Stanford FMTI biasing transparency toward corporate services
Key points:
- Stanford's FMTI claims to measure LLM transparency but actually measures how well-documented a commercial product is, not research transparency
- FMTI is systematically biased against openly released models — open-access research papers are excluded from evidence, yet BLOOM-Z still scored higher than closed APIs
- GPT-4 received 48% despite disclosing no non-trivial methodology details; a research-community index would give it 0%
- The scorecard approach encourages Goodhart's Law gaming: companies can hire staff to generate compliant documentation without substantive change
- FMTI conflates model artifacts (research outputs) with hosted model services (commercial products), asking research projects questions only relevant to SaaS businesses
- Model weights and training data release — most fundamental to research transparency — are worth only 3/100 points in the index
Connections: Eleutherai · Stanford Crfm · Openai · AI Transparency · AI Governance · Large Language Models