查看: 4333| 回复: 1
跳转到指定楼层
上一主题 下一主题
收起左侧

[找工就业] 统计类面试基本概念总结 (有需要进)

全局:

2020(7-9月)-Stat/Biostat硕士+1-3年 | 网上海投|大西雅图地区 分析|数据科学类全职@

注册一亩三分地论坛,查看更多干货!

您需要 登录 才可以下载或查看附件。没有帐号?注册账号

x


花了一些时间自己做了一些统计基本概念的总结,来应对面试可能会问到的问题。根据自己理解的角度做的总结,可能有理解和表达不精准的地方,可以一起探讨下
. Waral dи,. 1point3acres.com
有需要的同学可以当作一个参考:
.--
. 1point 3 acres

Sampling distribution/ sample mean/ how to compare the means of two samples
                i. Sample mean is an estimate of the 'true' mean for the population
                ii. Choose a p-value and run t-test and compare to threshold value to decide if we should reject Null hypothesis.

Confidence interval
                i. Confidence interval always comes with bootstrap sampling>10000, for each sampling we got a sample mean.. From 1point 3acres bbs
                ii. CI is a range, in which a certain percentage of the sample means fall in.
                iii. If the confidence level is 95%, it means 95% of the sample means fall into that range, which can be further explained that we are 95% sure that true population mean fall into that range.. 1point3acres.com
. check 1point3acres for more.
P-value
                i. P-value is a measure of statistical confidence. A measure of how confident you are that the model will still be valid if you have a lot more data.. From 1point 3acres bbs
                        a. The lower the p-value, the greater the statistical confidence.
                ii. P-value is used in the hypothesis testing when comparing two sample data and is the evidence against the null hypothesis.. 1point3acres
                        a. If the p-value is lower than the threshold value(0.05), we would like to reject the null hypothesis which supports two samples coming from the sample distribution.
                        b. There are only 5% of chance (too small) to get this extreme observation if the null hypothesis is true.
                        c. There are only 5% for the statistical test to be False Positive (too small) if the two samples do come from the same distribution (null hypothesis is true).
. Waral dи,
Multi-hypothesis testing
                i. The multiple hypothesis testing occurs when a number of individual hypothesis tests are considered simultaneously. . ----
                ii. In this case, the significance of individual tests no longer represents the combined sets.. 1point 3 acres

P-hacking
                i. P-hacking
                        a. In multi-hypothesis testing , P-hacking means the misuse of the analysis techniques so the results are misled by false positive.
                        b. Take the single result/observation to draw conclusion or adjust sample size during the experiment to achieve statistical significance.. 1point 3 acres
                        c. If the  two samples do come from the same distribution (null hypothesis is true), there are still 5% chance that the statistical tests will result in False Positive.
                ii. How to prevent P-hacking
                        a. False Discovery Rate method usually prevent P-hacking in multi-hypothesis testing by inputting every single p-value and output the adjusted p-value that is usually greater than the original ones.
                        b. Bonferroni Method (Reduce the threshold of P-value so we can reduce the false positive rate).-baidu 1point3acres

Power Analysis
                i. The power analysis determines what sample size will ensure the higher probability to correctly reject a null hypothesis
                ii. We can't adjust the sample size during the experiment because it will lead to P-hacking and only can do that before the experiment.
                iii. Effect size (metric to show the overlap of two distributions)/ desired power level / confident level

Central limit theorem
                i. No matter what distribution from which the sample means were calculated, the means of the collected samples (sample mean) are normally distributed
                ii. When doing experiment , we don't know the distribution where the group data come from, but CLT doesn't care since the sample means are always normally distributed
                iii. On top of that we conduct t-test and ANOVA to see If the means of two groups are statistically different.
                iv. Sample size > 30

Bias/Variance/Underfitting/Overfitting
                i. Bias (under-fitting) - The inability for a machine learning method to capture the true relationship-baidu 1point3acres
                ii. Variance (overfitting) - the difference of the performance in fitting among different  datasets.
                iii. The high variance means the hypothesis line fit the training set so well, but it did a terrible job in fitting the test set.
                iv. Overfitting - the function line fits the training set well but not the testing set.



大家喜欢的话支持以下哦,也帮助我可以看到更多有用的面经,谢谢~. Waral dи,

评分

参与人数 13大米 +37 收起 理由
czhmily + 1 给你点个赞!
nikilo + 1 赞一个
LLLXY + 1 很有用的信息!
s11012 + 2 很有用的信息!
pikado + 20 欢迎分享你知道的情况,会给更多积分奖励!

查看全部评分


上一篇:亚麻21Intern时间线求问
下一篇:准备明年回国,求各种面试准备,技巧和坑

本帖被以下淘专辑推荐:

  • · DS|主题: 224, 订阅: 39
🔗
yyunchien 2023-1-12 02:17:23 | 只看该作者
全局:
給你點讚!
回复

使用道具 举报

您需要登录后才可以回帖 登录 | 注册账号
隐私提醒:
  • ☑ 禁止发布广告,拉群,贴个人联系方式:找人请去🔗同学同事飞友,拉群请去🔗拉群结伴,广告请去🔗跳蚤市场,和 🔗租房广告|找室友
  • ☑ 论坛内容在发帖 30 分钟内可以编辑,过后则不能删帖。为防止被骚扰甚至人肉,不要公开留微信等联系方式,如有需求请以论坛私信方式发送。
  • ☑ 干货版块可免费使用 🔗超级匿名:面经(美国面经、中国面经、数科面经、PM面经),抖包袱(美国、中国)和录取汇报、定位选校版
  • ☑ 查阅全站 🔗各种匿名方法

本版积分规则

>
快速回复 返回顶部 返回列表