楼主: zhangjy529
跳转到指定楼层
上一主题 下一主题
收起左侧

DS 学习打卡贴

🔗
 楼主| zhangjy529 2017-2-5 01:40:36 | 只看该作者
全局:
Feb 3: 完成 COURSERA, using Python to access web data, 第二章: networks and sockets

评分

参与人数 1大米 +15 收起 理由
anonym + 15 坚持的不错,再接再厉!

查看全部评分

回复

使用道具 举报

🔗
白衣胜雪 2017-2-5 03:27:00 | 只看该作者
全局:
楼主,现在药厂还招人不
回复

使用道具 举报

🔗
 楼主| zhangjy529 2017-2-14 08:52:03 | 只看该作者
全局:
争取现在的一分一秒,才能改变现状, 为自己加油!
回复

使用道具 举报

🔗
 楼主| zhangjy529 2017-2-18 02:13:45 | 只看该作者
全局:
FEB 17: STATISTICS, multiple factor ANOVA summary: Independent t-test is used to compare the means of two independent population. when we need to compare the means of i(i>=3) independent samples, we use the analysis of ANOVA. For ANOVA, the data are regarded as random sample from K populations. ANOVA can also be used to see if there is a trend in the i groups and used to estimate the stnadard deviation. The extra sum of squares is used to do the F-TEST. We divide the total sum of variation into the error sum of squares and the sum of variation between groups. The multiple comparison is used to compare the means of two groups. For example, Bonferr method and Tukey method. And residual plot against the fitted value is used to check the constant variation assumption, normality plot of the residual plot is used to check the normality assumption. If data are collected in time, the run plot is used to check the independence assumption.

评分

参与人数 1大米 +15 收起 理由
anonym + 15 坚持的不错,再接再厉!

查看全部评分

回复

使用道具 举报

🔗
 楼主| zhangjy529 2017-2-19 10:13:36 | 只看该作者
全局:
FEB18: Using Pyhton to access webdata. 1. Regular expression. 2. Socket, request and response cycle 3. Urllib library does the socket for HTTP and makes web pages look like a file. Urllib does not return the head of a website. 4. Parse HTML with BeautifulSoup 5. Parse XML(eXtensible Markup Language) in Python: import xml-etree.ElementTree as ET. 6. JSON(JavaScript Object Notation). JSON is Javascript and looks like Python. JSON represents data as nested "lists" and "dictionaries"

评分

参与人数 1大米 +15 收起 理由
anonym + 15 坚持的不错,再接再厉!

查看全部评分

回复

使用道具 举报

🔗
 楼主| zhangjy529 2017-2-20 03:57:28 | 只看该作者
全局:
One-Way ANOVA Summary
For single factor experiments, the levels of the factors are called experiments. A completely randomized design randomized N=∑_(i=1)^a▒n_i  experiments to a treaments with ni units on the ith treatment. We assume that the observations y_(i,j) from the ith treatment are a random sample form N(μ_i,σ^2). And the samples from the a treatments are mutually independent.
The one-way ANOVA partitions the total sum of squares into two components: the treatment sum of squares(SSR) and the error sum of squares (SSE). The ratio of the sum of squares to its degree of freedom are called the mean square. The error mean squares (MSE) gives an unbiased estimate of σ^2 with v=N-a d.f. The overall null hypothesis is that the group mean are the same, and is tested using the F statistics . check 1point3acres for more.
F=MSA/MSE
Which has an F-DISTRIBUTION WITH A-1 DEGREE OF FREEOM and N-A D.F.
Multiple comparison methods are useful for identifying the treatments that differ from each other. Multiple tests can result in excessive false significance. Therefore we control the type one familywise error rate(FWE) of finding at least one false significant difference at a specified level α. The LSD method, which does pairwise two-sample t-tests each at level α does not control the FEW. The Bonferroni method controls the FEW conservationaly at level α by doing each pairwie t-test at level α/k, k=(a|2) is the number of pairwise comparisons. The Tukey method are exact when the sample sizes in each group are the same under the null hypothesis. Stepwise multiple testing are less conservative. . .и

评分

参与人数 1大米 +15 收起 理由
anonym + 15 坚持的不错,再接再厉!

查看全部评分

回复

使用道具 举报

🔗
 楼主| zhangjy529 2017-2-20 04:49:35 | 只看该作者
全局:
FEB 19: Command line interface commands, pwd, clear, ls, cd, mkdir, touch, cp, rm, mv, echo, date

评分

参与人数 1大米 +15 收起 理由
anonym + 15 坚持的不错,再接再厉!

查看全部评分

回复

使用道具 举报

🔗
 楼主| zhangjy529 2017-2-20 14:46:07 | 只看该作者
全局:
Feb19: MACHINE LEARNING, BIAS VARIANCE TRADE OFF, LINEAR REGRESSION,scikit-learn library

评分

参与人数 1大米 +15 收起 理由
anonym + 15 坚持的不错,再接再厉!

查看全部评分

回复

使用道具 举报

🔗
linbaobei001 2017-2-21 04:54:46 | 只看该作者
全局:
我看了看   我跟上面几个人的情况差不多,都是在职 转DS, 我打算建一个小组群,, 大家一起学习打卡   加我微信吧   备注 :DS复习     691290857
回复

使用道具 举报

🔗
 楼主| zhangjy529 2017-2-21 08:29:34 | 只看该作者
全局:
FEB 20: Machine learning with linear regression and logistic regression. (1) Check data with .head(), info(), etc (2) Exploratory data analysis (3) clean data, impute or drop missing values, change category variable to 0/1, e.g. sex. (4) divide into train and test (train_test_split) and fit, predict. (5) linear regression, check residual, and line plot of y_test vs prediction. (6) Logistic regression: confusion matrix or classification matrix

K nearest neighbours

评分

参与人数 1大米 +15 收起 理由
anonym + 15 坚持的不错,再接再厉!

查看全部评分

回复

使用道具 举报

您需要登录后才可以回帖 登录 | 注册账号
隐私提醒:
  • ☑ 禁止发布广告,拉群,贴个人联系方式:找人请去🔗同学同事飞友,拉群请去🔗拉群结伴,广告请去🔗跳蚤市场,和 🔗租房广告|找室友
  • ☑ 论坛内容在发帖 30 分钟内可以编辑,过后则不能删帖。为防止被骚扰甚至人肉,不要公开留微信等联系方式,如有需求请以论坛私信方式发送。
  • ☑ 干货版块可免费使用 🔗超级匿名:面经(美国面经、中国面经、数科面经、PM面经),抖包袱(美国、中国)和录取汇报、定位选校版
  • ☑ 查阅全站 🔗各种匿名方法

本版积分规则

>
快速回复 返回顶部 返回列表