📣 Back to School开学季 - VIP通行证5折优惠!蓝莓、Offer多多同步优惠
查看: 3403| 回复: 0
跳转到指定楼层
上一主题 下一主题
收起左侧

Nature文章,机器学习分析数据预测科学家发展潜力!~

全局:

注册一亩三分地论坛,查看更多干货!

您需要 登录 才可以下载或查看附件。没有帐号?注册账号

x

   文章挺有意思的……不知道被预测为发展潜力欠佳是神马感觉呢……  请各位看官高抬贵手赏几个大米啊!想看300分限制文章,绞尽脑汁赚大米


We research scientists often worry about the future of our careers. Is our research an exciting path or a dead end that will end our careers prematurely? Predicting scientific trajectories is a daily task for hiring committees, funding agencies and department heads who probe CVs searching for signs of scientific potential.
One popular measure of success is physicist Jorge Hirsch's h-index1, which captures the quality (citations) and quantity (number) of papers, thus representing scientific achievements better than either factor alone. A scientist has an h-index of n if he or she has published n articles receiving at least n citations each2. Einstein, Darwin and Feynman, for example, have impressive h-indices of 96, 63 and 53, respectively. According to Hirsch, anh-index of 12 for a physicist — meaning 12 papers with at least 12 citations each — could qualify him or her for tenure at a major university.
ILLUSTRATION BY DAVID PARKINS


However, the h-index3 and similar metrics4 can capture only past accomplishments, not future achievements5. Here we attempt to predict the future h-index of scientists on the basis of features found in most CVs.
We maintain that the best way of predicting a scientist's future success is for peers to evaluate scientific contributions and research depth, but think that our methods could be valuable complementary tools.
The typical research CV contains information on the number of publications, those in high-profile journals, theh-index and collaborators. One can also infer interdisciplinary breadth, the length and quality of training, the amount of funding received and even the standing of the scientist's PhD adviser. Such factors are taken into account for hiring decisions, but how should they be weighted? Fortunately, obtaining data on the scientific activities of individual researchers has never been easier. Using all of these features, we can begin to probe the scientific enterprise statistically.
Vital statistics
To construct a formula to predict future h-index, we assembled a large data set and analysed it using machine-learning techniques. Our initial sample from academictree.org — a crowd-sourced website listing scientists' mentors, trainees and collaborators — contains the names and institutions of about 34,800 neuroscientists, 2,000 scientists studying the fruitfly Drosophila and 1,300 evolutionary researchers. We matched these authors to records in Scopus, an online database of academic papers and citation data. We restricted our analysis to authors who had accrued an h-index greater than 4 (to exclude inactive scientists); to publications after 1995 (because electronic records are sparse before then); to authors who had published their first manuscript in the past 5–12 years; and to authors who were identifiable in Scopus.
That left us with 3,085 neuroscientists, 57 Drosophila researchers and 151 evolutionary scientists for whom we constructed a history of publication, citation and funding.
For each year since the first article published by a given scientist, we used the features that were available at the time to forecast their h-index a number of years into the future. For example, we reconstructed how the CV features of a scientist looked five years after publishing his or her first article, and found a relationship between those features and the reconstructed h-index five years on.
Starting with neuroscientists, we attempted to predict the h-index of each scientist 5 years ahead — a timescale relevant for tenure decisions — using a linear regression with elastic net regularization6(seeSupplementary Information). The model predicted the future h-index accurately, yielding a respectableR2=0.67, cross-validated across scientists (an R2 of 1 would imply that the model predicts the data perfectly). A simplified model containing only the number of published articles, the h-index, years since first publication, number of publications in prestigious neuroscience journals (Nature, Science, Nature Neuroscience, Neuron and the Proceedings of the National Academy of Sciences) and the number of distinct journals still performed nearly equally well (R2=0.66; see 'Predict your future h-index').
Box 1: Metrics: Predict your future h-index
These are approximate equations for predicting the h-index of neuroscientists in the future. They are probably reasonably precise for life scientists, but likely to be less meaningful for the other sciences. Try it for yourself online at go.nature.com/z4rroc.
Predicting next year (R2 = 0.92):
Predicting 5 years into the future (R2 = 0.67):
Predicting 10 years into the future (R2 = 0.48):
Key: n, number of articles written; h, current h-index; y, years since publishing first article; j, number of distinct journals published in; q, number of articles in Nature, Science, Nature Neuroscience, Proceedings of the National Academy of Sciences and Neuron.




Predicting the future careers of Drosophila and evolutionary scientists leads to somewhat worse predictions (R2=0.54 and R2=0.61, respectively, based on scientists 3–15 years into their careers) but still better than predictions based on the h-index alone (R2=0.38 and R2=0.39, respectively). This indicates that generalizations to other fields within and outside of life science may be limited1. But for neuroscientists, at least, the predictions extend well to longer periods of time, such as ten years into the future (R2=0.52). Over time, using just the h-index performs much worse than taking all features into account (see 'Paths to success', left panel).
The main five predictive features change in importance for predicting h-indices over increasingly longer periods (see 'Paths to success', right panel). The power of the h-index declines. The number of articles written, the diversity of publication in distinct journals and the number of articles published in five prestigious journals all become increasingly influential over time.
Future fortunes
It is risky to make any causal interpretations of these results. However, we will briefly speculate on why these features might be important predictors of future success. Some features directly affect the potential for a high h-index, such as the number of articles written. These features can also indirectly affect a scientist's future success, because scientists who are productive and publish many papers tend to remain productive. Publishing in many different journals may lead to fewer overlapping populations of scientists who cite the work, and hence higher growth potential for articles. A scientist who has published in several distinct journals is also likely to be someone with broad training who contributes in many ways. The number of publications in leading journals can increase the visibility of a scientist's other papers, past and future.
If promotion, hiring or funding were largely based on indices (h-index, the model used here or any other measure), then some scientists would adapt their behaviour to maximize their chances of success. Models such as ours that take into account several dimensions of scientific careers should be more difficult for researchers to game than those that focus on a single measure.
Our formula is particularly useful for funding agencies, peer reviewers and hiring committees who have to deal with vast numbers of applications and can give each only a cursory examination. Statistical techniques have the advantage of returning results instantaneously and in an unbiased way. Building and analysing massive data sets to track scientific careers could also help to identify potential gender, racial and other biases7, 8, 9and advance our understanding of how science develops.
Although our findings and predictions may not alleviate scientists' angst over their careers, the results offer some comfort by showing that the future is not so random. The occasional rejection of a paper may feel unjust and indiscriminate, but in the long run, such factors seem to average out, rendering h-index trajectories relatively predictable.

评分

参与人数 1大米 +20 收起 理由
ruly7170 + 20 感谢分享!

查看全部评分


上一篇:Coursera.org无法登陆的解决办法
下一篇:[Coursera] Stanford Machine Learning (Week #4)
您需要登录后才可以回帖 登录 | 注册账号
隐私提醒:
  • ☑ 禁止发布广告,拉群,贴个人联系方式:找人请去🔗同学同事飞友,拉群请去🔗拉群结伴,广告请去🔗跳蚤市场,和 🔗租房广告|找室友
  • ☑ 论坛内容在发帖 30 分钟内可以编辑,过后则不能删帖。为防止被骚扰甚至人肉,不要公开留微信等联系方式,如有需求请以论坛私信方式发送。
  • ☑ 干货版块可免费使用 🔗超级匿名:面经(美国面经、中国面经、数科面经、PM面经),抖包袱(美国、中国)和录取汇报、定位选校版
  • ☑ 查阅全站 🔗各种匿名方法

本版积分规则

>
快速回复 返回顶部 返回列表