楼主: 匿名
跳转到指定楼层
上一主题 下一主题
收起左侧

AB testing面试题请教

 
🔗
czhmily 2021-1-11 10:45:58 | 只看该作者
全局:
顶起 等大神们回复
回复

使用道具 举报

🔗
lsydd489 2021-1-11 11:31:16 | 只看该作者
全局:
啊这个我见过,叫early stopping,你可以搜一下,比如说你在50%的时候发现5%停下来,这时候你实际的alpha大约是8.6%,而不是5%。

评分

参与人数 1大米 +1 收起 理由
czhmily + 1 很有用的信息!

查看全部评分

回复

使用道具 举报

🔗
lsydd489 2021-1-11 11:48:12 | 只看该作者
全局:
这是我上udacity上的课的一段代码

这个函数不改参数的情况下返回的是你在半中间看一下如果满足p<5%就停止实验前后两个block的p value 以及 整个实验总体的p value
def peeking_sim(alpha = .05, p = .5, n_trials = 1000, n_blocks = 2, n_sims = 10000):    """. 1point3acres.com
    This function simulates the rate of Type I errors made if an early
    stopping decision is made based on a significant result when peeking ahead.

    Input parameters:
        alpha: Supposed Type I error rate
        p: Probability of individual trial success. Waral dи,
        n_trials: Number of trials in a full experiment
        n_blocks: Number of times data is looked at (including end)
        n_sims: Number of simulated experiments run

    Return:
        p_sig_any: Proportion of simulations significant at any check point,
        p_sig_each: Proportion of simulations significant at each check point
    """
    trials_per_block = np.ceil(n_trials / n_blocks).astype(int)
    data = np.random.binomial(trials_per_block, p, n_sims * n_blocks).reshape(n_sims, n_blocks)
. 1point 3 acres
      # standardize data
    data_cumsum = data.cumsum(axis = 1)
    block_sizes = trials_per_block * np.arange(1, n_blocks+1, 1)
    block_means = block_sizes * p
    block_sds   = np.sqrt(block_sizes * p * (1-p))
    data_zscores =(data_cumsum - block_means)/block_sds
. 1point 3acres
      # test outcomes
    z_crit = stats.norm.ppf(1-alpha/2)
    sig_flags = abs(data_zscores) > z_crit
    p_sig_any = (sig_flags.sum(axis = 1)>=1).mean()
    p_sig_each = sig_flags.mean(axis = 0)
. Χ
    return (p_sig_any, p_sig_each)


这段代码是求 corrected alpha:  
def peeking_correction(alpha = .05, p = .5, n_trials = 1000, n_blocks = 2, n_sims = 10000):. Waral dи,
    """-baidu 1point3acres
    This function uses simulations to estimate the individual error rate necessary
    to limit the Type I error rate, if an early stopping decision is made based on
    a significant result when peeking ahead.
    . 1point3acres.com
    Input parameters:
        alpha: Desired overall Type I error rate
        p: Probability of individual trial success
        n_trials: Number of trials in a full experiment.--
        n_blocks: Number of times data is looked at (including end)
        n_sims: Number of simulated experiments run
.1point3acres        
    Return:
        alpha_ind: Individual error rate required to achieve overall error rate
    """
   
    # generate data
    trials_per_block = np.ceil(n_trials / n_blocks).astype(int)
    data = np.random.binomial(trials_per_block, p, [n_sims, n_blocks])
   
    # standardize data
    data_cumsum = np.cumsum(data, axis = 1)
    block_sizes = trials_per_block * np.arange(1, n_blocks+1, 1)
    block_means = block_sizes * p.--
    block_sds   = np.sqrt(block_sizes * p * (1-p))
    data_zscores = (data_cumsum - block_means) / block_sds.--
   
    # find necessary individual error rate
    max_zscores = np.abs(data_zscores).max(axis = 1)
    z_crit_ind = np.percentile(max_zscores, 100 * (1 - alpha))
    alpha_ind = 2 * (1 - stats.norm.cdf(z_crit_ind))
    . Waral dи,
    return alpha_ind


回复

使用道具 举报

🔗
dyyyyyu 2021-1-11 13:17:32 | 只看该作者
全局:
czhmily 发表于 2021-1-11 04:09. 1point 3 acres
请问如果PM还是决定第一周停下的话,是不是可以调整p value ,例如以前都是p

前两天正好看到一篇medium post有相关的内容:https://medium.com/lime-eng/expe ... at-lime-bee846d62dd.google  и
这篇讲的是Lime怎么做experimentation,很详细的介绍了experimentation会涉及到的很多问题。里面在“Dealing with the Peeking Problem”那一段有提到他们如果被要求提前report结果(没有达到pre-power analysis要求的Sample size),只在confidence level达到99.9%的时候才claim stat sig。也讲了他们这样做的理由。供楼主参考一下。

评分

参与人数 3大米 +3 收起 理由
Cassie_LinHCCX + 1 很有用的信息!
sprezzatura + 1 给你点个赞!
czhmily + 1 很有用的信息!

查看全部评分

回复

使用道具 举报

🔗
Janeevans 2021-1-11 13:30:33 | 只看该作者
全局:
还有一种方法可以deal with this early stopping, 如果是要你设计AB experimentation system的话,可以用Bayesian method 去做动态调整,实现Optimal stopping,具体算法你可以google一下这几个关键词。

评分

参与人数 1大米 +1 收起 理由
czhmily + 1 很有用的信息!

查看全部评分

回复

使用道具 举报

🔗
czhmily 2021-1-11 14:48:23 | 只看该作者
全局:
dyyyyyu 发表于 2021-1-11 13:17
前两天正好看到一篇medium post有相关的内容:https://medium.com/lime-eng/experimentation-analysis-at ...

谢谢 没有花钱 看不到文章 能发个copy 或者帮总结一下吗
回复

使用道具 举报

🔗
dyyyyyu 2021-1-11 15:57:22 | 只看该作者
全局:
czhmily 发表于 2021-1-11 14:48
谢谢 没有花钱 看不到文章 能发个copy 或者帮总结一下吗
..
emmm copy了相关的段落:.--
. check 1point3acres for more.
Dealing with the Peeking Problem

Peeking (looking at test results before reaching the needed sample size (n)) causes the probability of a Type 1 error (rejecting a true null hypothesis) to increase. The most common way to deal with this is to calculate the required n before starting the test and then only making a decision when you have reached that sample size. However, the drawbacks to that solution are that:

1. We want to peek to stop bad things from happening or flawed experimental design >> Fix: we can still do this as long as we don’t report the interim results officially; we can just use it as a safeguard in case of negative things occurring.

2. We are too reliant on our power analysis calculations before the test starts (i.e. a priori effect), e.g. we overestimate the effect we expect to see or that we need to see from a business perspective >> Fix: we can always err on the side of a low effect size like 2% over 5%. (Means higher n necessary though so not very sustainable). Waral dи,

3. The p = .08 problem: It’s clear that we are almost at significance and directionally moving that way. But we need to let the experiment run a bit more, collect more n to get to the 5% threshold >> Fix: we can always err on the side of requiring a larger n than we think we need. But again this doesn’t seem like the optimal solution.

At Lime, our solution is to run a normal power analysis before the experiment starts to calculate the day on which we should have sufficient sample size to check results. We can only check the final results one time on that planned day and need a 95% confidence level to claim the results are statistically significant. If we want to check results before that time (or are requested to do so), then we can only claim results are stat sig if they meet a 99.9% confidence level as opposed to the initial 95%. This reduces the likelihood that we are falsely rejecting the null hypothesis by recalculating the impact multiple times. We decided that this solution is better suited than using Bayes Factors (another common solution in the industry) since it is a simpler calculation and we are still building out our internal experimentation analysis tools.
(我觉得这篇文章全文都很不错,而且很长。。。你浏览器开个incognito应该就可以看了)

评分

参与人数 1大米 +1 收起 理由
czhmily + 1 很有用的信息!

查看全部评分

回复

使用道具 举报

🔗
redeye1 2021-1-11 23:44:04 | 只看该作者
本楼:
全局:
mark 一下
回复

使用道具 举报

🔗
icesample 2021-1-12 00:34:15 | 只看该作者
全局:
ZeZheng 发表于 2021-1-11 03:18
一般做ab test都是根据想要达到的statistical power,一般是0.8,来计算一个sample size,然后用sample siz ...
..
你好,我明白你说的点。我的想法是,如果一周时间p value <5%, 这个时候sample size 只是一半,type 2 error 高,但是type I error 低(?),reject null 是跟type I error 相关的啊。换句话说,我们在一半sample size (statistical power 较低)的情况下就能reject null 了,在variance 不变增加sample size,岂不是更有可能reject null 了吗?
回复

使用道具 举报

🔗
wujiayikelly 2021-1-12 02:50:36 | 只看该作者
全局:
不能,如果你需要10天跑完,每一天的alpha是0.05,10天下来总的false positive rate = 1-(1-0.05)^10,远远大于0.05

看书里说一个比较形象的example是篮球比赛,总共需要打四场才能判断胜负,不能说打完一场我们队赢了就结束吧 :)

评分

参与人数 1大米 +1 收起 理由
czhmily + 1 很有用的信息!

查看全部评分

回复

使用道具 举报

您需要登录后才可以回帖 登录 | 注册账号
隐私提醒:
  • ☑ 禁止发布广告,拉群,贴个人联系方式:找人请去🔗同学同事飞友,拉群请去🔗拉群结伴,广告请去🔗跳蚤市场,和 🔗租房广告|找室友
  • ☑ 论坛内容在发帖 30 分钟内可以编辑,过后则不能删帖。为防止被骚扰甚至人肉,不要公开留微信等联系方式,如有需求请以论坛私信方式发送。
  • ☑ 干货版块可免费使用 🔗超级匿名:面经(美国面经、中国面经、数科面经、PM面经),抖包袱(美国、中国)和录取汇报、定位选校版
  • ☑ 查阅全站 🔗各种匿名方法

本版积分规则

>
快速回复 返回顶部 返回列表