The Big Data Brain Drain: Why Science is in Trouble
data-scienceacademiaopen-sourcereproducibilityscientific-softwarecareer
Abstraction: Academia's publish-or-perish model drives data-skilled scientists to industry
Key points:
- Skills for successful scientific research (statistics, computing, algorithm-building, software design) are now identical to industry skills, causing researcher drain from academia
- "Publish-or-perish" model discourages software development: time spent on open, reproducible code is time not spent writing papers, the primary academic currency
- Buckheit and Donoho: "An article about computational science is not the scholarship itself — it is merely advertising"; actual scholarship is the complete software and instructions
- LHC produces 10GB/s; LSST will produce 15TB/night — modern science entirely dependent on sophisticated data software
- NIH base postdoc salary under $40,000/year ($50k after 7 years); same skills command several times that in industry first-year roles
- Proposed fixes: reproducibility requirements in publication, new tenure criteria including open software, dedicated software-focused academic tracks, higher postdoc pay
Connections: Jake Vanderplas · Scikit Learn · Data Science · Reproducibility · Open Source Software
Source: https://jakevdp.github.io/blog/2013/10/26/big-data-brain-drain/