查看: 12851| 回复: 66
跳转到指定楼层
上一主题 下一主题
收起左侧

DS 学习打卡贴

全局:

注册一亩三分地论坛,查看更多干货!

您需要 登录 才可以下载或查看附件。没有帐号?注册账号

x
本人现在在药厂做 SAS PROGRAMMER. 想转行做DS. 所以决定每天学习。 希望每天都能打卡, 大家相互监督, 早点找到好的DS工作。. 1point3acres.com
. 1point 3acres
统计: 打算先复习基础概念。 教材: STATISTICS AND DATA ANALYSIS.

编程: 打算先学习PYTHON. COURSA: Programming for Everybody (Getting Started with Python).

JAN 22:
统计: 计划完成STATISTICS AND DATA ANALYSIS. CHAPTER 5: SAMPLING DISTRIBUTION OF STATISTICS, CHAPTER 6: BASIC CONCEPTS OF INFERENCES
编程: 完成PYTHON, 第一节和第二节课


. Χ


补充内容 (2017-1-27 15:09):
JAN 26: 修完COURSERA PYTHON 第一门课.
SAS: 通过UCLA 网站复习了如何用SAS 做 t-test, ANOVA, chi-square test, binomial test, Wilcoxon-Mann-Whitney test,Kruskal Wallis test,Wilcoxon signed rank sum te...

补充内容 (2017-1-28 01:54):
Jan 27: 终极目标: GOOGLE STATISTICIAN/Quantitative Analyst. 希望半年后的今天, 能坐在GOOGLE 的食堂里喝咖啡。fighting!
. .и
补充内容 (2017-1-29 13:20):. 1point3acres
Jan 27: SAS: 复习PROC REPORT. 重新了解了COLUMN, DEFINE, COMPUTE, BREAK, 以及通过ODS 把 把结果输出到PDF, HTML
JAN 28: PYTHON: 完成 PYTHON DATA STRUCTURE, STRINGS, FILES, LISTS 三章
-baidu 1point3acres
补充内容 (2017-1-30 04:55):
JAN 29:  教材: STATISTICS AND DATA ANALYSIS. 完成第五章: sampling distribution of sample mean, variance, student t-distribution, Fisher's F-Distribution, central limit theorem, order statistics

补充内容 (2017-2-1 01:53):
JAN 30: PYTHON: 完成 PYTHON DATA STRUCTURE, TUPLES章节, 修完这门课了。现在还想修另外三门PYTHON课程,但是学费好贵呀。 肉疼

补充内容 (2017-2-2 14:55):-baidu 1point3acres
Feb 1: Python: 完成 using Python to access web data, 第一章: regular expressions

补充内容 (2017-2-5 01:40):
Feb 3: 完成 using Python to access web data, 第二章: networks and sockets

补充内容 (2017-2-5 04:27):
Feb 4, 统计: 完成STATISTICS AND DATA ANALYSIS. CHAPTER 6: BASIC CONCEPTS OF INFERENCES, Point Estimation, Confidence interval estimation, hypothesis testing
.1point3acres
补充内容 (2017-2-5 15:12):
Feb 3: 完成 using Python to access web data, networks and sockets, programs that surf the web, web service and XML, JSON and the REST Architecture

补充内容 (2017-2-6 10:53):
FEB 5: 新增学习计划: COURSERA: DATA SCIENCE 10 COURSE SPECIALIZATION
COURSERA: Managing Big Data with MySQL
COURSERA: STANDARD MACHINE LEARNING

补充内容 (2017-2-6 10:54):.
新增学习计划: STATISTICS: REGRESSION, LOGISTIC REGRESSION, MULTIVARIATE STATISTICS, ANOVA, A/B TESTING, Categorical data analysis

补充内容 (2017-2-8 02:57):
Feb 7, 统计: STATISTICS AND DATA ANALYSIS. C 7:Inferences for single sample, inference for mean (large sample, Z distribution), inference for mean(small sample, t-distribution), and variance(chi-s...

补充内容 (2017-2-9 08:10):
Feb 8, 统计: STATISTICS AND DATA ANALYSIS. C 8:Inferences for two sample, independent samples and matched samples, compare means of two population

补充内容 (2017-2-9 14:17):
FEB 8: PYTHON, 完成COURSERA, USING DATABASE WITH PYTHON

补充内容 (2017-2-11 11:31):
Feb 10, 统计: STATISTICS AND DATA ANALYSIS. C 9:Inferences on proportions, compare two proportions, one-way count data
. 1point 3acres
补充内容 (2017-2-12 11:02):
Feb 10, 统计: STATISTICS AND DATA ANALYSIS. C 10: Simple linear regression and regression diagonistic. 1point3acres.com

补充内容 (2017-2-13 02:18):
FEB 12: 任务, 统计完成multiple regression, one-way ANOVA, PYTHON: 总结PYTHON 的五门课程, 继续UDEMY MACHINE LEARNING 课程

补充内容 (2017-2-13 15:43):.--
FEB 12: UDEMY: PYTHON for data science and machine learning, NUMPY, PANDAS
. 1point3acres
补充内容 (2017-2-16 15:31):
FEB 15: UDEMY: PYTHON for data science and machine learning,  matplotlib, seaborn, plotly, cufflinks. check 1point3acres for more.
Statistics: Finish One way ANOVA

补充内容 (2017-2-16 15:32):
FEB 15: UDEMY: PYTHON for data science and machine learning,  matplotlib, seaborn, plotly, cufflinks. Χ
Statistics: Finish One way ANOVA. ----

补充内容 (2017-2-16 15:32):. 1point3acres.com
FEB 15: UDEMY: PYTHON for data science and machine learning,  matplotlib, seaborn, plotly, cufflinks
Statistics: Finish One way ANOVA. .и

补充内容 (2017-2-18 01:57):. From 1point 3acres bbs
FEB 17: STATISTICS, FINISH half of multiple factor ANOVA. From 1point 3acres bbs

补充内容 (2017-2-21 09:35):
FEB 20: LINEAR REGRESSION, LOGISTRIC REGRESSION AND K Nearest Neighbor by Python

补充内容 (2017-2-22 14:01):
FEB 21: PCA and RECOMMENDER SYSTEM

补充内容 (2017-2-24 05:40):
FEB 23: NATURAL LANGUAGE PROCESSING. NLTK librarypipeline=Pipeline([. .и
  CountVectorizer(analyzer=text_process)),  # strings to token integer counts
    ('tfidf', TfidfTransformer()), MultinomialNB

补充内容 (2017-4-2 02:19):
Apr 1, 2017:  从最开始准备找data scientist 工作到现在三个月, 有一些电话面试和onsite. 总结的经验是准备不充分。但也知道应该怎么准备。首先概率论和基础统计,工作写SQL, 完成ANDREW machine learning课程. 1point 3acres

补充内容 (2017-4-8 01:35):
4月7号: 学习计划, 完成review of prob, collecting data, summarize and explore data, sampling distribution of statistics, basic concepts of inference. check 1point3acres for more.

补充内容 (2017-4-14 01:26):
4.7-4.13: 完成review of prob, collecting data, summarize and explore data, sampling distribution of statistics,

补充内容 (2017-4-16 09:52):. check 1point3acres for more.
4.15 finish inference of one sample, two sample, proportions and count data, non-linear regression

补充内容 (2017-4-17 07:52):
4.16: Simple linear regression, probability model for SLR, least square fit, goodness of fit by R-square and ANOVA, inference for intercept and slope, prediction of future observation, regression dia

补充内容 (2017-4-27 10:23):
4.18-4.26, 复习完应用统计基础第二遍, Andrew machine learning, neural network, SVM, cluster, variance and bias, recommender system, train/validation/test.

补充内容 (2017-5-3 06:36):
这两周面试了两个dream company 的phone interview, 都挂了。博士基础没打牢固再加转行,真的不是两三个月就能补起来的。 打起精神, 再接再厉,希望8月份能找到好公司的data scientsit

补充内容 (2017-5-11 02:19):
5.10   下周电话面试。现在在把PYTHON, machine learning重新复习一遍。 机会来了就要抓住, 加油。

补充内容 (2017-6-2 09:06):
June 1, 有幸拿到dream company 的onsite。 还有三周就onsite. 现在工作也很忙。考验我们时候到了,提高学习和工作效率, 加油!. check 1point3acres for more.
-baidu 1point3acres
补充内容 (2017-6-3 06:41):
June 2,继续打卡。计划这个周末完成 regression analysis教材。

补充内容 (2017-6-16 06:59):
Jun15, 最近工作任务很多, 每天加班到九点。 不管怎么忙, 复习计划一点不能落下。 这两天看了logistic regression, 感觉思路一下子清楚了很多。
.google  и
补充内容 (2017-10-27 10:19):
前面几个月被工作搞死了。 现在终于忙完了。有过一些面试, 大部分是小公司。 发现自己编程达不到他们的要求,hands on 能力不够。 都挂了。 现在父母过来帮忙, 赶紧回来打卡。 希望明年年初能找到工作。 加油!

补充内容 (2018-8-2 02:20):
注册了一个KAGGLE 账户做MACHINE LEARNING PROJECT。大家如果觉得还不错帮忙点一下UP。 如果有什么问题也希望留言交流。我会及时回复。谢谢大家!
https://www.kaggle.com/junyingzhang2018/kernels

评分

参与人数 2大米 +160 收起 理由
DL + 10 感谢分享!
anonym + 150 坚持的不错,再接再厉!

查看全部评分


上一篇:不知道大家这两个项目会怎么选:USF Analytics & UT MSBA
下一篇:請問有人了解USC MS in Analytics?
推荐
 楼主| zhangjy529 2017-3-6 06:46:47 | 只看该作者
全局:
March 5: Logistic regression.
The log odds ratio, called the logit, has the linear relationship with X. The change rate is beta.phi(x)(1-phi(x)). At phi(x)=P(Y=1|X)=1/2, the change rate is the biggest. And x=-alpha/beta. exp(beta) is an odds ratio, the odds at X=x+1 divided by the odds at X=x.
1. Logistic regression with retrospective studies: For sample of subjects haveing Y=1 (cases) and Y=0 (controls), the value of X is observed. Evidence exists of an association if the distribution of X values differs between cases and controls.
2. Type of inferences. (1) wald statistis, z=beta/SE . Under H0, Z^2 is approximately chi-square(1) distribution.  (2). The likelihood ratio test uses twice the difference between the maximized log ikelihood at beta_hat and at beta=0. and also has approximately chi-square(1) . (3) The score test uses the log likeliho at beta=0 through the derivative of the log likelihood (the score function) at that point. The test ompare the sufficient statistics for beta to its null expected value.
confidence interval BY WALD method
3. Checking goodness of fit: For any type of binary data, one way to detect lack of fit uses the likelihood-ratio test to compare the model to more complext model. A more complex model might contain non-linear effect or interaction terms. If the more complex model do not fit better, this provides some assurance that the model chosen is reasonable. If the explanatory variable is category, we can compare the observed counts with the fitted values using a Pearson chi-sqaure or likelihood ratio G-SQUARE. It is chi-square distribution with DF equal to the number of parameters - the number of parameters in the model.. Waral dи,
4. logit models with categorical predictors (factors). Use dummy variables in logit model. Linear logit model for IX2 tables. . Χ
5. Multiple logistic regression. The parameter beta-i refers to the effect of x_i on the log odds that Y=1, controlling the other x_j. exp(beta_j) is the multiplcative effect on the odds of 1-unit increase in x_i, at fixed levels of other x_j. -baidu 1point3acres
6. Goodness of fit as a likelihood-ratio test: -2(L_0-L_1) tests whether certain model parameters are zero by comparing the log likelihood L_1 for the fitted model M-1 with L_0 for a simpler model M_0. If M_0=M and M_1 is the saturated model. In test whether M fits, we test whether all parameters in the saturated model but not in M equal zero. The df is the difference in the number of parameters in the model.
7. Fitting the multiple logistic regression by MLE. Newton-Raphson iterative method for logistic regession.

评分

参与人数 1大米 +15 收起 理由
anonym + 15 坚持的不错,再接再厉!

查看全部评分

回复

使用道具 举报

推荐
 楼主| zhangjy529 2017-2-20 03:57:28 | 只看该作者
全局:
One-Way ANOVA Summary
For single factor experiments, the levels of the factors are called experiments. A completely randomized design randomized N=∑_(i=1)^a▒n_i  experiments to a treaments with ni units on the ith treatment. We assume that the observations y_(i,j) from the ith treatment are a random sample form N(μ_i,σ^2). And the samples from the a treatments are mutually independent.
The one-way ANOVA partitions the total sum of squares into two components: the treatment sum of squares(SSR) and the error sum of squares (SSE). The ratio of the sum of squares to its degree of freedom are called the mean square. The error mean squares (MSE) gives an unbiased estimate of σ^2 with v=N-a d.f. The overall null hypothesis is that the group mean are the same, and is tested using the F statistics
F=MSA/MSE
Which has an F-DISTRIBUTION WITH A-1 DEGREE OF FREEOM and N-A D.F..
Multiple comparison methods are useful for identifying the treatments that differ from each other. Multiple tests can result in excessive false significance. Therefore we control the type one familywise error rate(FWE) of finding at least one false significant difference at a specified level α. The LSD method, which does pairwise two-sample t-tests each at level α does not control the FEW. The Bonferroni method controls the FEW conservationaly at level α by doing each pairwie t-test at level α/k, k=(a|2) is the number of pairwise comparisons. The Tukey method are exact when the sample sizes in each group are the same under the null hypothesis. Stepwise multiple testing are less conservative.

评分

参与人数 1大米 +15 收起 理由
anonym + 15 坚持的不错,再接再厉!

查看全部评分

回复

使用道具 举报

推荐
 楼主| zhangjy529 2017-2-18 02:13:45 | 只看该作者
全局:
FEB 17: STATISTICS, multiple factor ANOVA summary: Independent t-test is used to compare the means of two independent population. when we need to compare the means of i(i>=3) independent samples, we use the analysis of ANOVA. For ANOVA, the data are regarded as random sample from K populations. ANOVA can also be used to see if there is a trend in the i groups and used to estimate the stnadard deviation. The extra sum of squares is used to do the F-TEST. We divide the total sum of variation into the error sum of squares and the sum of variation between groups. The multiple comparison is used to compare the means of two groups. For example, Bonferr method and Tukey method. And residual plot against the fitted value is used to check the constant variation assumption, normality plot of the residual plot is used to check the normality assumption. If data are collected in time, the run plot is used to check the independence assumption.

评分

参与人数 1大米 +15 收起 理由
anonym + 15 坚持的不错,再接再厉!

查看全部评分

回复

使用道具 举报

🔗
aKisspy 2017-1-23 13:41:58 | 只看该作者
全局:
new grad找工作!同打卡。已经刷完了绿皮的brain teaser和coursera 的machine learning。现在在搞R打project。下一步刷cpp和esl那本书
回复

使用道具 举报

🔗
 楼主| zhangjy529 2017-1-25 08:27:24 | 只看该作者
全局:
根据总结其他人的经验, 需要具备以下能力。   Ttest, Regression, ANOVA, Logistic Regression, DOE, Machine Learning, Data Mining, MapReduce, SQL, R/Matlab, Python, Java
回复

使用道具 举报

🔗
msesmart 2017-1-26 00:19:12 | 只看该作者
全局:
同在职跳槽,刚刷完programming for everybody 前两门,想找本书复习统计基础知识。lz可以发个statistics and data analysis 书的链接吗?我们可以一起刷
回复

使用道具 举报

🔗
 楼主| zhangjy529 2017-1-27 01:22:14 | 只看该作者
全局:
Jan 26, 2017: 编程: 完成PYTHON, 已经完成第六周课程。 再有一节课就能毕业拿证书了。 学习这门课程给我很大的信息。 老师讲的特别好, 生动形象, 简单易懂。 给我这种无编程基础的人很大信息。 准备听完这门课以后再把后续的四门课修完。 加油。

评分

参与人数 1大米 +15 收起 理由
anonym + 15 坚持的不错,再接再厉!

查看全部评分

回复

使用道具 举报

🔗
 楼主| zhangjy529 2017-1-27 05:24:05 | 只看该作者
全局:
msesmart 发表于 2017-1-26 00:19
同在职跳槽,刚刷完programming for everybody 前两门,想找本书复习统计基础知识。lz可以发个statistics a ...

我用的是这本。研究生的一门基础统计课的教材。 我因为统计知识感觉都忘得差不多了, 所以选择从最简单的开始。 打算快速看完这本在读其他人推荐的。  
https://www.amazon.com/Statistics-Data-Analysis-Elementary-Intermediate/dp/0137444265
回复

使用道具 举报

🔗
msesmart 2017-1-27 05:29:29 | 只看该作者
全局:
zhangjy529 发表于 2017-1-26 16:24
我用的是这本。研究生的一门基础统计课的教材。 我因为统计知识感觉都忘得差不多了, 所以选择从最简单 ...

Lz你有这本书的电子版吗?求发邮箱yy3ja@virginia.edu谢啦
回复

使用道具 举报

🔗
 楼主| zhangjy529 2017-1-27 15:01:39 | 只看该作者
全局:
msesmart 发表于 2017-1-27 05:29
Lz你有这本书的电子版吗?求发邮箱谢啦

这个我是以前国内买的影印版, 没有电子书。 我觉得你找本类似的就行了。
回复

使用道具 举报

🔗
seeker2013 2017-1-28 07:02:35 | 只看该作者
全局:
Same here! Analyst II in healthcare company now willing to jump into other industry! very interested in DS  
回复

使用道具 举报

🔗
 楼主| zhangjy529 2017-1-30 04:58:35 | 只看该作者
全局:
seeker2013 发表于 2017-1-28 07:02
Same here! Analyst II in healthcare company now willing to jump into other industry! very interested ...

加油! 感觉现在在药厂编程统计知识用的很少。 编程很多也是用的MACRO. 对自己的成长不是很好。转DS学的东西会多一些, 统计知识,编程能力, 解决问题的能力也能提高很多, 更具挑战性。
回复

使用道具 举报

您需要登录后才可以回帖 登录 | 注册账号
隐私提醒:
  • ☑ 禁止发布广告,拉群,贴个人联系方式:找人请去🔗同学同事飞友,拉群请去🔗拉群结伴,广告请去🔗跳蚤市场,和 🔗租房广告|找室友
  • ☑ 论坛内容在发帖 30 分钟内可以编辑,过后则不能删帖。为防止被骚扰甚至人肉,不要公开留微信等联系方式,如有需求请以论坛私信方式发送。
  • ☑ 干货版块可免费使用 🔗超级匿名:面经(美国面经、中国面经、数科面经、PM面经),抖包袱(美国、中国)和录取汇报、定位选校版
  • ☑ 查阅全站 🔗各种匿名方法

本版积分规则

>
快速回复 返回顶部 返回列表