楼主: modifiedname
跳转到指定楼层
上一主题 下一主题
收起左侧

Data Scientist 炼成记录-更新完毕2018年12月 | 机器学习练成记录 - 已开新帖

   
🔗
god91425 2015-11-18 14:48:49 | 只看该作者
全局:
Mark...大赞一亩三分地
回复

使用道具 举报

🔗
jamie1987321 2015-11-28 21:48:47 | 只看该作者
本楼:
全局:
大牛啊!
回复

使用道具 举报

🔗
chaoyue2500 2015-12-15 14:24:44 | 只看该作者
全局:
很有帮助的帖子,谢谢
回复

使用道具 举报

🔗
lzhang06 2015-12-18 15:59:08 | 只看该作者
全局:
回复马克一下
回复

使用道具 举报

全局:
mark----------------------------------------
回复

使用道具 举报

🔗
 楼主| modifiedname 2016-1-7 09:10:09 | 只看该作者
全局:
Some update:
spent the last couple of months focusing on python (general) programming. . 1point 3 acres
scipy/numpy/pandas and general programming are sufficiently different -- knowing the former won't help you on coding interviews at all
个人用R做分析比用py多一点,不过也没有强烈preference, both have nice tools/vis, R is still better for "deep sh*t" in stats, Py merges with other parts of the system easily.
. From 1point 3acres bbs
Finally understand enough about web dev to work on backend DS or hacking projects  - hands on project is the only way to learn.
also spent some time in code structure, tests, intermediate topics

still can't 刷题 in leetcode in any decent manner -- coding面一直强烈短板,依靠其他方面的deep expertise掩盖过关。 ..
java/scala skills are still dismal,至今没用上过

lost some interest in research type, deeper stats
became more interested in solving real life problems with data and code, the data part may or may not involve deep stats
strategic thinking - what is the best approach in terms of cost and benefit and maintainability? 80-20 rules: MVP and iterate - which method in my toolset can solve the problem with the best tradeoff of effort and reward? when is a good time to go the complex route?

I think it's a mentality change - solving real life problems with technology, vs playing with deep/cool stuff in tech that happens to solve some problems as a side effect

Also got more interested in system design as a result -- sexy machine learning algo/model selection and validation plays a very small role in practical systems -- not that it's not important to get things right in principle (it still has to be right), but having the right system: extract and store features (feature engineering), efficiently use memory, predict/classify in a timely manner, etc are probably 95-98% of the work, except for at very large companies with a large team of PhDs, which celebrate 0.1-1% improvements.
回复

使用道具 举报

🔗
verachen 2016-1-8 12:38:17 | 只看该作者
本楼:
全局:
mark~~~~
回复

使用道具 举报

🔗
ccx-fdu 2016-1-12 10:26:34 | 只看该作者
全局:
新人报到一下
回复

使用道具 举报

🔗
ccx-fdu 2016-1-12 11:44:20 | 只看该作者
全局:
我现在是在加拿大读Bioresource Engineering的,主要做modeling,但以后想找data analysist类型的工作,然后已经开始自学Data science了。现在主要用Python来刷leetcode的题,还停留在easy水平。SQL自学完了,现在在coursera上不断刷machine learning和algorithm的课程,有点迷茫,各位能指点一下应该怎么学么?感觉小白一头扎进来到处都摸不清。而且因为我专业跟statistic或者CS都没什么关系,找Data方向工作的时候会成为劣势和阻碍么?谢谢各位帮忙啦!

点评

actually sounds good. Pick up some domain knowledge (maybe look at PM interviews for some ideas)  发表于 2016-1-12 11:48
回复

使用道具 举报

🔗
ccx-fdu 2016-1-12 12:18:03 | 只看该作者
全局:
ccx-fdu 发表于 2016-1-12 11:44
.. 我现在是在加拿大读Bioresource Engineering的,主要做modeling,但以后想找data analysist类型的工作,然 ...

So glad to have your kind respond! I will try to do better. THX~
回复

使用道具 举报

您需要登录后才可以回帖 登录 | 注册账号
隐私提醒:
  • ☑ 禁止发布广告,拉群,贴个人联系方式:找人请去🔗同学同事飞友,拉群请去🔗拉群结伴,广告请去🔗跳蚤市场,和 🔗租房广告|找室友
  • ☑ 论坛内容在发帖 30 分钟内可以编辑,过后则不能删帖。为防止被骚扰甚至人肉,不要公开留微信等联系方式,如有需求请以论坛私信方式发送。
  • ☑ 干货版块可免费使用 🔗超级匿名:面经(美国面经、中国面经、数科面经、PM面经),抖包袱(美国、中国)和录取汇报、定位选校版
  • ☑ 查阅全站 🔗各种匿名方法

本版积分规则

>
快速回复 返回顶部 返回列表