楼主: zhangjy529
跳转到指定楼层
上一主题 下一主题
收起左侧

DS 学习打卡贴

🔗
 楼主| zhangjy529 2017-3-6 06:46:47 | 只看该作者
全局:
March 5: Logistic regression.
The log odds ratio, called the logit, has the linear relationship with X. The change rate is beta.phi(x)(1-phi(x)). At phi(x)=P(Y=1|X)=1/2, the change rate is the biggest. And x=-alpha/beta. exp(beta) is an odds ratio, the odds at X=x+1 divided by the odds at X=x.
1. Logistic regression with retrospective studies: For sample of subjects haveing Y=1 (cases) and Y=0 (controls), the value of X is observed. Evidence exists of an association if the distribution of X values differs between cases and controls.
2. Type of inferences. (1) wald statistis, z=beta/SE . Under H0, Z^2 is approximately chi-square(1) distribution.  (2). The likelihood ratio test uses twice the difference between the maximized log ikelihood at beta_hat and at beta=0. and also has approximately chi-square(1) . (3) The score test uses the log likeliho at beta=0 through the derivative of the log likelihood (the score function) at that point. The test ompare the sufficient statistics for beta to its null expected value.
confidence interval BY WALD method
3. Checking goodness of fit: For any type of binary data, one way to detect lack of fit uses the likelihood-ratio test to compare the model to more complext model. A more complex model might contain non-linear effect or interaction terms. If the more complex model do not fit better, this provides some assurance that the model chosen is reasonable. If the explanatory variable is category, we can compare the observed counts with the fitted values using a Pearson chi-sqaure or likelihood ratio G-SQUARE. It is chi-square distribution with DF equal to the number of parameters - the number of parameters in the model.
4. logit models with categorical predictors (factors). Use dummy variables in logit model. Linear logit model for IX2 tables.
5. Multiple logistic regression. The parameter beta-i refers to the effect of x_i on the log odds that Y=1, controlling the other x_j. exp(beta_j) is the multiplcative effect on the odds of 1-unit increase in x_i, at fixed levels of other x_j.
6. Goodness of fit as a likelihood-ratio test: -2(L_0-L_1) tests whether certain model parameters are zero by comparing the log likelihood L_1 for the fitted model M-1 with L_0 for a simpler model M_0. If M_0=M and M_1 is the saturated model. In test whether M fits, we test whether all parameters in the saturated model but not in M equal zero. The df is the difference in the number of parameters in the model.
7. Fitting the multiple logistic regression by MLE. Newton-Raphson iterative method for logistic regession.

评分

参与人数 1大米 +15 收起 理由
anonym + 15 坚持的不错,再接再厉!

查看全部评分

回复

使用道具 举报

🔗
白沂秋 2017-3-7 03:52:03 | 只看该作者
全局:
楼主的打卡贴好励志啊,我也是刚开始工作,暂时做IT consultant 想转DS,跟着楼主一起学!
回复

使用道具 举报

🔗
 楼主| zhangjy529 2017-3-7 12:05:17 | 只看该作者
全局:
白沂秋 发表于 2017-3-7 03:52
楼主的打卡贴好励志啊,我也是刚开始工作,暂时做IT consultant 想转DS,跟着楼主一起学!

现在感觉好多东西要学呀,加油
回复

使用道具 举报

🔗
 楼主| zhangjy529 2017-3-17 13:23:53 | 只看该作者
全局:
March 16, 这周跟几个公司的recruiter聊了一下。 希望下周会有电话面试。简历投的比较仓促, 现在觉得有很多都没有准备到。 加油!加油!

评分

参与人数 1大米 +15 收起 理由
anonym + 15 坚持的不错,再接再厉!

查看全部评分

回复

使用道具 举报

🔗
 楼主| zhangjy529 2017-3-18 08:45:39 | 只看该作者
全局:
March 17: 过来报道一下, 今天开始复习machine learning的理论指示。 在下周一以前完成。 同时晚上KAGGLE上面的项目和问卷调查。

评分

参与人数 1大米 +15 收起 理由
anonym + 15 坚持的不错,再接再厉!

查看全部评分

回复

使用道具 举报

🔗
 楼主| zhangjy529 2017-3-20 10:22:06 | 只看该作者
全局:
March 19, Decision tree, bagging (bootstrap of aggregation), random forest, boosting

评分

参与人数 1大米 +15 收起 理由
anonym + 15 坚持的不错,再接再厉!

查看全部评分

回复

使用道具 举报

🔗
qingff 2017-3-29 22:23:10 | 只看该作者
全局:
MARK! 向楼主学习!
回复

使用道具 举报

🔗
Love_data 2017-3-31 09:03:39 | 只看该作者
全局:
linbaobei001 发表于 2017-2-22 00:22
你说的绿皮的 brain teaser 能给个链接吗 或者大概什么样的 找不到呢,,,
. Waral dи,
同求。。。
回复

使用道具 举报

🔗
 楼主| zhangjy529 2017-4-2 02:20:52 | 只看该作者
全局:
Apr 1, 2017:  从最开始准备找data scientist 工作到现在三个月, 有一些电话面试和onsite. 总结的经验是准备不充分。但也知道应该怎么准备。首先概率论和基础统计,工作写SQL, 完成ANDREW machine learning课程。 虽然后面两周没有面试,正好把基础知识赶紧不起来。准备三周以后继续投简历。加油!!!

评分

参与人数 1大米 +15 收起 理由
anonym + 15 坚持的不错,再接再厉!

查看全部评分

回复

使用道具 举报

🔗
 楼主| zhangjy529 2017-4-2 02:32:10 | 只看该作者
全局:
Aril 1, 2017. 学习大纲。 1. 复习statistics教材 2. Probability教材 3. 看SAS PROC SQL, 工作中用SQL写程序 3. 完成ANDREW machine learning 教材 4. Prepare past projects (paper and code) and can explain to people easily. 5. 完成udacity A|B testing course. 6. finish COURSERA Bayes statistics course. 7. Do machine learning projects on Kaggle.com using Python and R. 8. Review regression, ANOVA, analysis of covariance, general linear model and generalized linear model

评分

参与人数 1大米 +15 收起 理由
anonym + 15 坚持的不错,再接再厉!

查看全部评分

回复

使用道具 举报

您需要登录后才可以回帖 登录 | 注册账号
隐私提醒:
  • ☑ 禁止发布广告,拉群,贴个人联系方式:找人请去🔗同学同事飞友,拉群请去🔗拉群结伴,广告请去🔗跳蚤市场,和 🔗租房广告|找室友
  • ☑ 论坛内容在发帖 30 分钟内可以编辑,过后则不能删帖。为防止被骚扰甚至人肉,不要公开留微信等联系方式,如有需求请以论坛私信方式发送。
  • ☑ 干货版块可免费使用 🔗超级匿名:面经(美国面经、中国面经、数科面经、PM面经),抖包袱(美国、中国)和录取汇报、定位选校版
  • ☑ 查阅全站 🔗各种匿名方法

本版积分规则

>
快速回复 返回顶部 返回列表