查看: 3928| 回复: 11
跳转到指定楼层
上一主题 下一主题
收起左侧

note for data jobs

 
全局:

注册一亩三分地论坛,查看更多干货!

您需要 登录 才可以下载或查看附件。没有帐号?注册账号

x
本帖最后由 期待阳光 于 2023-8-21 22:00 编辑

  • p-value
.--
When conducting a hypothesis test, the extremeness of the test statistic in relation to the expected distribution under H0 plays a crucial role. Let's break this concept down:
1. The Distribution under H0
When the null hypothesis H0 is true, we expect our test statistic to follow a certain distribution. This distribution is essentially a representation of all possible test statistics we might compute if H0 were indeed the reality, and we could sample from our population infinitely many times.
2. Observed Test Statistic:. check 1point3acres for more.
When we conduct an experiment or a study, we collect data and compute a test statistic from this data. This single value is our observed metric, which we want to evaluate in the context of H0.
3. Comparing to the Distribution:
Given the distribution under H0, we can determine how typical or atypical our observed test statistic is. If it's a value that's common under H0, then our observed results are consistent with the null hypothesis. If it's atypical (or "extreme"), it suggests that our data is not easily explained by the null hypothesis.
4. The p-value as a Measure of Extremeness:
The p-value quantifies this extremeness. If you think of the distribution of the test statistic under H0 as a continuum:
  • A small p-value (often < 0.05) indicates that the observed test statistic is in the tails of the distribution (rare/extreme under H0).
  • A large p-value indicates that the observed test statistic is near the center of the distribution (common/expected under H0​).

5. Decision Making:
Based on the p-value, researchers decide whether to reject H0H0​ in favor of the alternative hypothesis HaHa​:
  • Small p-value: Typically leads to rejecting H0H0​. The data we observed is unlikely to have occurred by random chance alone, given H0.
  • Large p-value: We do not reject H0H0​. Our data is consistent with what we might expect if H0H0​ were true.

Conclusion:
The entire process of hypothesis testing revolves around this concept of "extremeness." It helps researchers differentiate between results that might have just occurred by random variability and those that might suggest a genuine effect or difference.

  • how to map test statistic to p-value
    . 1point 3 acres



t-distribution: The test statistic in a t-test follows the t-distribution. The specific shape of this distribution depends on the degrees of freedom, often related to sample size.
  • Locate t-statistic: Suppose you've computed a t-statistic of, say, 2.3 from your data. You'd then consider where 2.3 falls on the t-distribution with the relevant degrees of freedom.
  • Determine Tail Area:

    • Two-tailed test: If you're conducting a two-tailed t-test, you'd find the area in the right tail beyond 2.3 and double it (because you're considering both tails).
    • One-tailed test: If it's a one-tailed test (where you're testing if the mean of one group is greater than the other), you'd just find the area in the right tail beyond 2.3.
  • . From 1point 3acres bbs
    • After determining the p-value, you compare it to your significance level (often 0.05) to decide whether to reject the null hypothesis..--
. 1point3acres
. 1point 3acres
  • what is p-value relationship with the probability of type one error?

1. Definition of p-value:
The p-value represents the probability of observing a test statistic as extreme as, or more extreme than, the one calculated from the sample data, assuming that the null hypothesis (H0) is true.
2. Definition of Type I error:
A Type I error, also known as a false positive or alpha error, occurs when we incorrectly reject a true null hypothesis. The probability of committing a Type I error is denoted by α (alpha).
.
Key Takeaway:
The p-value is a measure derived from the data, indicating the strength of the evidence against H0H0​. In contrast, αα is a predefined threshold, usually set by the researcher or by convention, representing the maximum allowable probability of committing a Type I error. The decision to reject H0H0​ or not is based on comparing the p-value to α.

  • what is the power of a test
  • Type II Error (False Negative): Before diving into power, it's essential to understand the concept of a Type II error. A Type II error occurs when a test fails to reject the null hypothesis even though the alternative hypothesis is true. It's often denoted by the symbol β.
  • Power = 1 - β: The power of a test is the complement of the probability of making a Type II error. Hence, Power = 1 - β.

. ----
  • an example for Power Support for Decision Making: Especially in applied fields, decisions might be based on statistical tests. Knowledge of a test's power provides decision-makers with an additional metric to gauge the reliability of the findings.


. .и
### Scenario: Testing a New Drug


**Background:**  
A pharmaceutical company has developed a new drug intended to lower blood pressure. Before releasing it to the market, they need to conduct clinical trials to test its efficacy.


**Study Design:**  . 1point3acres.com
The company conducts a randomized controlled trial where participants are assigned to one of two groups:

. 1point3acres
1. Treatment group: Receives the new drug.
2. Control group: Receives a placebo.
.
.google  и
The aim is to test if there's a significant difference in blood pressure reduction between the two groups after a set period.


### Initial Results:
. Χ
.--
The researchers collect data and run an appropriate statistical test. They find a p-value of 0.06, just slightly above the common 0.05 threshold for significance. Based on this p-value alone, they might decide not to reject the null hypothesis, suggesting that the drug doesn't have a statistically significant effect on blood pressure reduction.

. ----
### Power Analysis:. ----


Before making a final decision, the company's analysts decide to consider the power of the test. They compute that, given the sample size and observed effect size, the test had a power of only 0.55 (or 55%) to detect a significant difference at the 0.05 level.


### Decision Making:


Armed with this knowledge, the decision-makers in the company have a more nuanced perspective:


1. The p-value was near the significance threshold, suggesting a possible effect.
2. The test had only 55% power, which means there was a relatively high chance of making a Type II error (not detecting an effect when one exists).


Given the potential financial and health implications, the company might decide to conduct a larger study to increase the power, hoping to achieve a clearer result. They wouldn't want to discard a potentially effective drug based on an underpowered study, especially when the observed p-value was close to the significance threshold.. ----


### Conclusion:
. Waral dи,
. .и
In this example, the concept of power informed decision-making by providing an additional lens to evaluate the reliability of the study findings. Instead of solely relying on the p-value, decision-makers used power to weigh the risks of Type II errors, leading to a more informed choice about the next steps in the drug development process.

评分

参与人数 5大米 +7 收起 理由
Taromochi + 1 给你点个赞!
小小小呀小爻 + 1 给你点个赞!
律__ + 1 给你点个赞!
colospring + 1 赞一个
DL + 3 给你点个赞!

查看全部评分


上一篇:Data Science Take Home challenge书
下一篇:请问有人了解American Credit Acceptance的面试流程吗 (DA岗)
推荐
 楼主| 期待阳光 2023-9-24 08:05:46 | 只看该作者
全局:
Time series ..
Understanding the reasons behind the calculation of ACF (Autocorrelation Function) and PACF (Partial Autocorrelation Function) is crucial for interpreting their patterns, especially in the context of time series modeling. Let's dive deeper:

### ACF (Autocorrelation Function):

1. **What it represents**: ACF measures the correlation between a series and its lagged values.

2. **Why calculate it**:
   - To identify the relationship or dependency between current and past values in a time series.
. 1point3acres.com    - To determine the appropriate lag length for models like ARIMA.
   - To distinguish between AR and MA processes. A tailing off ACF and a sharp cut-off in PACF indicate an AR process, whereas the reverse suggests an MA process.

3. **Calculation**:
   - For a time series \(y_t\), the autocorrelation at lag \(k\) is:
     \[ \rho_k = \frac{Cov(y_t, y_{t-k})}{Var(y_t)} \]. 1point 3acres
   - This formula finds the correlation between a series and its lagged version. Essentially, it's seeing how today's value relates to the value from k days ago.. 1point 3acres
. From 1point 3acres bbs
### PACF (Partial Autocorrelation Function):

1. **What it represents**: PACF measures the correlation between a series and its lagged values, but after removing the variations already explained by the intervening time series values.

2. **Why calculate it**:
   - To determine the actual correlation between the residual (which remains after removing influences of other intervening observations) of the current series and its lagged values.
   - Like ACF, PACF helps determine the appropriate lag length for ARIMA models.
   - PACF helps to determine the order of an AR process.

3. **Calculation**:
   - The partial autocorrelation at lag \(k\) is the correlation that results after removing the effect of any correlations due to the terms at shorter lags.
   - Formally, it's determined by regressing the data on its own lags and using the coefficients. For instance, to find the PACF at lag 2, we regress \(y_t\) on \(y_{t-1}\) and \(y_{t-2}\).
   - The coefficient of \(y_{t-2}\) is the PACF at lag 2. It represents the direct effect of \(y_{t-2}\) on \(y_t\).

In essence, ACF and PACF are tools that provide insights into the time dependency structure of a series, which is foundational for modeling exercises like ARIMA. By understanding their patterns and knowing their computational background, one can make better-informed decisions when modeling time series data.

Example :
Let's walk through a concrete example using a hypothetical AR(1) time series.

Imagine we have a simple AR(1) model:. 1point 3 acres

\[ y_t = 0.6y_{t-1} + \epsilon_t \]

Where:
- \( y_t \) is the value at time \( t \).
- \( \epsilon_t \) is white noise.
- The coefficient is 0.6, indicating the impact of the previous term on the current term.

Let's generate 10 observations for this series:

Time (t) | \( y_t \)
--- | ---
1 | 5. Χ
2 | 5.8 (5 + 0.6(5) + random noise)
3 | 6.1
4 | 5.5
5 | 5.9
6 | 6.2
7 | 5.8
8 | 5.7
9 | 6.0
10 | 6.3

(Note: The above series is hypothetical and is generated based on the AR(1) model with some randomness added.)

### ACF:
..
The ACF will measure the correlation between \( y_t \) and its lags. Given the AR(1) nature of our series:

- Lag 1 ACF will be close to 0.6 (the coefficient).
- Lag 2 ACF will be \( 0.6^2 \) = 0.36.
- Lag 3 ACF will be \( 0.6^3 \) = 0.216.
... and so on. The ACF will decay exponentially.. Waral dи,
. Χ
### PACF:

For an AR(1) model:
- PACF at lag 1 will be close to the AR(1) coefficient (0.6 in our case).
- PACF for lags 2 and beyond will be close to zero.

Why? Because the PACF only captures the correlation for a specific lag, excluding the effect of lags before it. Since our series is AR(1), only the first lag matters.
. ----
### Conclusion:. 1point 3 acres

Based on the hypothetical series and the ACF and PACF patterns:
- ACF decays exponentially, indicating the series is an AR process.
- PACF has a significant spike only at lag 1 and then is close to zero, pointing towards an AR(1) process.

This example illustrates the behavior of an AR(1) process. In practice, real-world data will have noise, and the ACF and PACF plots won't be this clean. However, the general patterns should be recognizable.

  • MA and AR yt mean and vairance. ----

let's dive into the mean and variance for the MA(1) process. Let's say our MA(1) process is defined by the equation:

\[ y_t = \epsilon_t + \theta_1 \epsilon_{t-1} \]

Where:
- \( y_t \) is the time series value at time \( t \).
- \( \epsilon_t \) is a white noise error term at time \( t \) with zero mean and variance \( \sigma^2 \).
- \( \theta_1 \) is the coefficient for the MA(1) process.
.
### Mean of \( y_t \):

E(\( y_t \)) = E(\( \epsilon_t + \theta_1 \epsilon_{t-1} \))
. 1point3acres.com
Since the expected value of a white noise process is zero, \( E(\epsilon_t) = 0 \) and \( E(\epsilon_{t-1}) = 0 \).

So, E(\( y_t \)) = 0 + 0 = 0

The mean of \( y_t \) is zero for the MA(1) process.
. Waral dи,
### Variance of \( y_t \):

Var(\( y_t \)) = Var(\( \epsilon_t + \theta_1 \epsilon_{t-1} \))

Using properties of variance:

Var(\( y_t \)) = Var(\( \epsilon_t \)) + \(\theta_1^2\) Var(\( \epsilon_{t-1} \)) + 2\(\theta_1\)Cov(\( \epsilon_t, \epsilon_{t-1} \)). From 1point 3acres bbs

Given that \( \epsilon_t \) and \( \epsilon_{t-1} \) are white noise, they are uncorrelated. Therefore, their covariance, Cov(\( \epsilon_t, \epsilon_{t-1} \)), is zero.

Also, Var(\( \epsilon_t \)) = \( \sigma^2 \) and Var(\( \epsilon_{t-1} \)) = \( \sigma^2 \).

Plugging these values in, we get:

Var(\( y_t \)) = \( \sigma^2 \) + \(\theta_1^2 \sigma^2\). ----

So, the variance of the MA(1) process is \( \sigma^2 \) (1 + \(\theta_1^2\)).

To sum up:
- The mean of \( y_t \) in an MA(1) process is zero..1point3acres
- The variance is \( \sigma^2 \) (1 + \(\theta_1^2\)), where \( \sigma^2 \) is the variance of the white noise process and \( \theta_1 \) is the MA(1) coefficient.

let's break down the mean and variance for an AR(2) model. An AR(2) model can be represented as:

\[ y_t = c + \phi_1 y_{t-1} + \phi_2 y_{t-2} + \varepsilon_t \]. 1point3acres

Where:. 1point3acres.com
- \( c \) is a constant.
- \( \phi_1 \) and \( \phi_2 \) are the AR(2) coefficients.
- \( \varepsilon_t \) is white noise with mean 0 and variance \( \sigma^2 \).

### Mean of \( y_t \) for AR(2):
The expected value (or mean) of \( y_t \) is:

\[ E(y_t) = E(c + \phi_1 y_{t-1} + \phi_2 y_{t-2} + \varepsilon_t) \]

Using linearity of expectation and given that \( E(\varepsilon_t) = 0 \), we get:

\[ E(y_t) = c + \phi_1 E(y_{t-1}) + \phi_2 E(y_{t-2}) \]
. 1point 3 acres
If the process is stationary, \( E(y_t) \) will be the same for all \( t \), so let's call it \( \mu \):

\[ \mu = c + \phi_1 \mu + \phi_2 \mu \]

Rearranging and combining like terms:

\[ \mu (1 - \phi_1 - \phi_2) = c \].google  и

From which:
. 1point 3acres
\[ \mu = \frac{c}{1 - \phi_1 - \phi_2} \].1point3acres

### Variance of \( y_t \) for AR(2):
. 1point 3 acres
The variance of \( y_t \) is:

\[ Var(y_t) = Var(c + \phi_1 y_{t-1} + \phi_2 y_{t-2} + \varepsilon_t) \]

Given the variance of a constant is 0, and assuming \( y_{t-1} \), \( y_{t-2} \), and \( \varepsilon_t \) are uncorrelated, we have:

\[ Var(y_t) = \phi_1^2 Var(y_{t-1}) + \phi_2^2 Var(y_{t-2}) + \sigma^2 \]
..
For a stationary process, the variance of \( y_t \) will be constant for all \( t \), so we'll call it \( \gamma_0 \):

\[ \gamma_0 = \phi_1^2 \gamma_0 + \phi_2^2 \gamma_0 + \sigma^2 \]

From which, combining like terms:

\[ \gamma_0 (1 - \phi_1^2 - \phi_2^2) = \sigma^2 \]

Given the stationarity conditions of the AR(2) model, the solution for \( \gamma_0 \) would be:
. Waral dи,
\[ \gamma_0 = \frac{\sigma^2}{1 - \phi_1^2 - \phi_2^2} \]

This is a basic understanding of the mean and variance for an AR(2) model. The stationarity conditions for an AR(2) model are more involved than for AR(1). In general, the roots of the characteristic polynomial associated with the AR(2) process must lie outside the unit circle for the process to be stationary.
回复

使用道具 举报

推荐
 楼主| 期待阳光 2023-9-17 07:01:34 | 只看该作者
全局:
Maximize likelihood estimation ..

Maximum likelihood estimation (MLE) is a method of estimating the parameters of a statistical model. Given some observed data and a statistical model, the method provides estimates for the model's parameters that maximize the likelihood of the observed data given those parameters.

The likelihood function \( L(\theta; X) \) quantifies how well the data \( X \) is described by the model with parameters \( \theta \). The maximum likelihood estimates are the values of \( \theta \) that maximize this function. Mathematically, this is expressed as:

\[
\hat{\theta}_{MLE} = \arg\max_{\theta} L(\theta; X)
\]

In other words, we are looking for the values of \( \theta \) that make the observed data \( X \) most probable.

To find these values, one typically takes the natural logarithm of the likelihood function to create the log-likelihood function \( \log L(\theta; X) \). This is often easier to work with mathematically, and it turns products into sums, making the calculations more tractable.

\[
\hat{\theta}_{MLE} = \arg\max_{\theta} \log L(\theta; X)
\]

To find the values of \( \theta \) that maximize the log-likelihood, you typically take the derivative of the log-likelihood function with respect to the parameters, set the derivatives equal to zero, and solve for the parameters. Depending on the complexity of the model and the likelihood function, finding the maximum likelihood estimates could involve straightforward algebraic solutions, numerical methods, or even advanced optimization techniques.

MLE for linear regression

Maximum Likelihood Estimation (MLE) can be used to find the best-fitting line in linear regression. In the simplest case, let's assume we have a linear model defined as:

\[
y = \beta_0 + \beta_1 x + \epsilon
\]

where \( y \) is the dependent variable, \( x \) is the independent variable, \( \beta_0 \) and \( \beta_1 \) are the parameters we want to estimate, and \( \epsilon \) is the error term, usually assumed to be normally distributed with mean zero and variance \( \sigma^2 \).

The likelihood function for a given set of data \( (x_1, y_1), (x_2, y_2), \ldots, (x_n, y_n) \) would be the joint probability of observing all these \( y_i \)'s given the \( x_i \)'s and parameters \( \beta_0 \), \( \beta_1 \), and \( \sigma^2 \). Assuming the errors \( \epsilon \) are independent and normally distributed, the likelihood function \( L \) is:

\[.1point3acres
L(\beta_0, \beta_1, \sigma^2) = \prod_{i=1}^{n} \frac{1}{\sqrt{2\pi\sigma^2}} e^{-\frac{(y_i - \beta_0 - \beta_1 x_i)^2}{2\sigma^2}}.
\]

Taking the natural logarithm, we get the log-likelihood function \( \log L \):

\[
\log L = -\frac{n}{2} \log(2\pi) - \frac{n}{2} \log(\sigma^2) - \frac{1}{2\sigma^2} \sum_{i=1}^{n} (y_i - \beta_0 - \beta_1 x_i)^2
\]

To maximize \( \log L \), we need to find \( \beta_0 \), \( \beta_1 \), and \( \sigma^2 \) that maximize this function. In other words, we take the partial derivatives with respect to \( \beta_0 \), \( \beta_1 \), and \( \sigma^2 \), set them equal to zero, and solve for these parameters.

In practice, this will yield the same estimates for \( \beta_0 \) and \( \beta_1 \) as the least squares method in the context of simple linear regression. This is because maximizing the log-likelihood for a linear regression model with normally distributed errors is mathematically equivalent to minimizing the sum of the squared errors, which is the objective function in least squares estimation.
.1point3acres
Nonetheless, the approach through MLE has the benefit of generalizing more easily to other kinds of distributions for the error term \( \epsilon \), should that be necessary.

Simple concrete example
Let's consider a simple concrete example with a small dataset.
. From 1point 3acres bbs
**Dataset**:
```
x: 1, 2, 3, 4, 5
y: 2.1, 4.2, 6.0, 7.9, 10.1
```. Waral dи,
.
We want to fit a linear regression model \( y = \beta_0 + \beta_1 x \) to this data.

**Step 1**: Define the likelihood function.
.
Given the assumption that the errors are normally distributed, the likelihood function for observing the given y-values given the x-values and parameters \( \beta_0 \) and \( \beta_1 \) is:

\[
L(\beta_0, \beta_1, \sigma^2) = \prod_{i=1}^{5} \frac{1}{\sqrt{2\pi\sigma^2}} e^{-\frac{(y_i - \beta_0 - \beta_1 x_i)^2}{2\sigma^2}}
\]. Waral dи,

**Step 2**: Define the log-likelihood function.

Taking the natural logarithm:

\[
\log L = -\frac{5}{2} \log(2\pi) - \frac{5}{2} \log(\sigma^2) - \frac{1}{2\sigma^2} \sum_{i=1}^{5} (y_i - \beta_0 - \beta_1 x_i)^2. 1point 3 acres
\]

**Step 3**: Maximize the log-likelihood.

To keep this explanation tractable, I'll summarize this step. In practice, you would take the partial derivatives of \( \log L \) with respect to \( \beta_0 \), \( \beta_1 \), and \( \sigma^2 \), set them to zero, and solve the resulting equations. This will yield values for \( \beta_0 \), \( \beta_1 \), and \( \sigma^2 \) that maximize the likelihood of observing the data.

For simple linear regression, this approach will yield the same values for \( \beta_0 \) and \( \beta_1 \) as the least squares method:

\[
\beta_1 = \frac{n(\Sigma xy) - (\Sigma x)(\Sigma y)}{n\Sigma x^2 - (\Sigma x)^2}.1point3acres
\]
\[
\beta_0 = \frac{\Sigma y - \beta_1 (\Sigma x)}{n}
\]
. .и
Computing for the given data:

\[
\beta_1 = \frac{5(110.7) - (15)(30.3)}{5(55) - (15)^2} = 2.02
\]
\[
\beta_0 = \frac{30.3 - 2.02(15)}{5} = 0.08
\]

So, the linear regression line that best fits our data (in the least squares sense and under our assumptions) is:. 1point 3 acres

\[. 1point3acres
y = 0.08 + 2.02x
\]

Again, using MLE with normal error assumptions for simple linear regression leads to the same results as using the least squares method. The power of MLE is more evident when moving beyond these simple scenarios.
回复

使用道具 举报

推荐
 楼主| 期待阳光 2023-8-22 10:31:15 | 只看该作者
全局:
  • p-value to non-technnical person:there is a probability of less than 5% that the result could have happened by chance.
  • Central Limit Theorem (CLT) and null hypothosis:
The concept of the sampling distribution under (the null hypothesis) is often related to the Central Limit Theorem (CLT) in many statistical tests. Let's explore this connection:

Central Limit Theorem (CLT):
The CLT states that, given a sufficiently large sample size, the sampling distribution of the sample mean (or sum) will be approximately normally distributed, regardless of the population distribution, provided the population has a finite variance. This approximation to normality becomes better as the sample size increases.

. 1point3acres.com Sampling Distribution under H0:
When conducting hypothesis tests, we want to understand the behavior of our test statistic under the assumption that the null hypothesis is true. The theoretical distribution of our test statistic, if we were to repeatedly sample from the population while H0 is true, is called the sampling distribution under H0.

Connection with CLT:. From 1point 3acres bbs
For many tests, particularly those involving means like the t-test, the CLT plays a crucial role:

t-tests: When we compare sample means, the CLT tells us that, given a sufficiently large sample, the distribution of sample means will be approximately normal. This normal distribution becomes the basis for the sampling distribution under H0. If sample sizes are smaller and we're comparing means, the t-test uses the t-distribution, which is a family of distributions that approaches the normal distribution as sample size increases.. Χ
Proportion tests: When testing proportions, the sampling distribution of the sample proportion (given a large enough sample size) will also be approximately normal due to the CLT.
Other tests: Even for tests that don't directly involve means or sums, understanding the distribution of the test statistic under H0 is vital. In some cases, the distribution may not be normal, and the CLT might not be directly applicable. For instance, the chi-square test has its test statistic following a chi-square distribution.
In many situations, the CLT provides the theoretical foundation for understanding the behavior of our test statistic under the null hypothesis, but it's essential to recognize when and how it applies.

  • correlation and covariance

Both correlation and covariance are statistical measures that describe the relationship between two variables. Here's a breakdown of their meanings:

1. **Covariance**:. Χ
   - **Definition**: Covariance measures the degree to which two variables change together. If the variables increase together, the covariance is positive. If one variable tends to increase while the other decreases, the covariance is negative.
   
   - **Formula**:
     \[. Waral dи,
     \text{Cov}(X, Y) = \frac{\sum_{i=1}^{n} (X_i - \bar{X})(Y_i - \bar{Y})}{n}
     \]. check 1point3acres for more.
     where:
     - \(X_i\) and \(Y_i\) are data points,
     - \(\bar{X}\) and \(\bar{Y}\) are the means of the variables X and Y, respectively,
     - \(n\) is the number of data points.
     
   - **Interpretation**: The sign of the covariance gives a direction of the relationship between the variables. However, its magnitude is not easy to interpret as it's not normalized and depends on the units of the variables.

2. **Correlation**:

   - **Definition**: Correlation measures the strength and direction of a linear relationship between two variables. It is standardized, meaning it always takes a value between -1 (perfect negative linear relationship) and 1 (perfect positive linear relationship), with 0 indicating no linear relationship.
   
   - **Formula**: The most commonly used form of correlation is the Pearson correlation coefficient, represented as \( r \):
     \[
     r = \frac{\sum_{i=1}^{n} (X_i - \bar{X})(Y_i - \bar{Y})}{\sqrt{\sum_{i=1}^{n} (X_i - \bar{X})^2 \sum_{i=1}^{n} (Y_i - \bar{Y})^2}}
     \]. .и
     
   - **Interpretation**: The sign of \( r \) indicates the direction of the relationship (positive or negative). The magnitude, which is always between -1 and 1, indicates the strength of the relationship. A value close to 1 or -1 indicates a strong linear relationship, while a value close to 0 indicates a weak linear relationship.


Key Difference:

The main difference between the two lies in their units and interpretability. Covariance is not standardized, so its value can range from negative infinity to positive infinity, making it somewhat challenging to interpret on its own. In contrast, correlation is dimensionless and standardized, offering a clear interpretation of the strength and direction of a linear relationship.

In summary, while both correlation and covariance give information about the relationship between two variables, correlation offers a normalized and more interpretable measure of the strength and direction of a linear relationship.
回复

使用道具 举报

🔗
 楼主| 期待阳光 2023-8-22 10:43:30 | 只看该作者
全局:
  • one-sample test:

A one-sample test, as the name suggests, involves testing a single sample of data against a known standard or theoretical expectation. It's used to determine if there is statistical evidence that the associated population parameter (e.g., mean) significantly differs from a hypothesized value. Here are the key points:
..
    - **One-Sample t-test for the Mean**: This is used when the population variance is unknown. The goal is to test if the population mean (\( \mu \)) is different from a specified value (\( \mu_0 \)).

    - **One-Sample z-test for the Mean**: This is used when the population variance is known. It also tests if the population mean is different from a specified value, but it's less commonly applied in practice because the population variance is rarely known.. 1point3acres

    - **Binomial Test**: This is used for binary or categoric


In summary, a one-sample test compares the mean or proportion of a single sample to a known standard or specified value to infer if the observed data provides evidence to suggest a significant difference from the hypothesized parameter.

  • paired sample test vs two-sample test.

Both the paired sample (or dependent) test and the two-sample (or independent) test compare means between two groups, but the context and the structure of the groups differ. Here's a breakdown:

1. **Paired Sample Test (Dependent Sample Test)**:
. check 1point3acres for more.
    - **Context**: Used when you're comparing two sets of related or matched data. This often means that the two sets of data come from the same group, measured at two different times or under two different conditions. Examples include before-and-after studies or comparing the same individuals under two different treatments.
. check 1point3acres for more.
    - **Hypotheses**:
      - \( H_0 \): \( \mu_{d} = 0 \) (The mean difference between the paired observations is zero.)
      - \( H_a \): \( \mu_{d} \neq 0 \) (The mean difference between the paired observations is not zero.).--
. 1point3acres.com
    - **Test Statistic**:
      The paired t-test uses the differences between the paired observations. Then, the mean and standard deviation of these differences are used to determine the test statistic.
. .и
    - **Assumptions**:
      - Observations are independent.
      - Differences between pairs should be approximately normally distributed, especially important for small sample sizes.

2. **Two-Sample Test (Independent Sample Test)**: ..

    - **Context**: Used when you have two separate, independent groups to compare. This could mean comparing males vs. females, treatment vs. control groups, etc., where individuals in one group have no direct relation to individuals in the other group.

    - **Hypotheses**:
      - \( H_0 \): \( \mu_{1} = \mu_{2} \) (The population means of the two groups are equal.)
      - \( H_a \): \( \mu_{1} \neq \mu_{2} \) (The population means of the two groups are not equal.). 1point 3acres
. 1point 3 acres
    - **Test Statistic**:
      The independent t-test compares the means of the two groups, considering both the sample sizes and variances. Variations of this test consider whether the two groups have equal variances (pooled t-test) or different variances (Welch's t-test).

    - **Assumptions**:
      - Observations within each group are independent.
      - Data in each group should be approximately normally distributed.
      - Variance of data in each group is equal (for pooled t-test). This assumption is relaxed in Welch's t-test.

**Key Difference**:

The central distinction is in the nature of the data. In the paired sample test, the two sets of data are related, often representing "before" and "after" measurements or two measurements under different conditions of the same subjects. In the two-sample test, the data sets come from distinct, independent groups. The paired test uses the differences between paired measurements, while the two-sample test compares the means of two separate groups directly.
回复

使用道具 举报

🔗
 楼主| 期待阳光 2023-8-25 11:16:50 | 只看该作者
全局:
  • Standard deviation and standard error

are both measures of dispersion or variability, but they are used for different purposes and in different contexts. Below is a comparison:

### Standard Deviation (\(s\))

1. **Definition**: Standard deviation is a measure of the amount of variation or dispersion in a set of values. It gives an idea of how spread out the values are around the mean.

2. **Formula**:
   - **Sample**: \( s = \sqrt{\frac{\sum{(x_i - \bar{x})^2}}{n-1}} \)
   - **Population**: \( \sigma = \sqrt{\frac{\sum{(x_i - \mu)^2}}{N}} \). check 1point3acres for more.

3. **Application**:
    - Used to quantify the amount of variation in a data set.
    - Helpful in identifying outliers or understanding data distribution.

4. **Units**: Same as the units of the data set.

5. **Interpretation**: A larger standard deviation indicates a wider dispersion around the mean, whereas a smaller standard deviation indicates that the values are closer to the mean..1point3acres

6. **Independent of Sample Size**: Standard deviation is not generally affected by the size of the sample.

### Standard Error (\(SE\))

1. **Definition**: Standard error is a measure of how much the sample mean \( \bar{x} \) is expected to vary from the true population mean \( \mu \).

2. **Formula**:
   - **Standard Error of the Mean**: \( SE = \frac{s}{\sqrt{n}} \)
   - \(s\) is the sample standard deviation and \(n\) is the sample size.

3. **Application**: .google  и
    - Used in hypothesis testing and confidence intervals..--
    - Useful in gauging the accuracy of \( \bar{x} \) as an estimate of \( \mu \).

4. **Units**: Same as the units of the data set..

5. **Interpretation**: A smaller standard error indicates that the sample mean is a more accurate reflection of the actual population mean.

6. **Dependent on Sample Size**: Standard error decreases as the sample size increases, given that the sample statistic is a more accurate estimate of the population parameter for larger samples.

### Summary

- **Standard deviation** is used to measure variability within a single sample..
- **Standard error** is used to measure the variability of a sample statistic (like the mean) across multiple samples from the same population.

In practice, standard deviation is often used for descriptive statistics to provide an understanding of the data distribution, while standard error is usually used for inferential statistics when drawing conclusions about the population from a sample.
回复

使用道具 举报

🔗
 楼主| 期待阳光 2023-8-28 00:51:14 | 只看该作者
全局:
  • Analysis of Variance (ANOVA)
Analysis of Variance (ANOVA) is a statistical method used to analyze differences among group means in a sample. ANOVA is based on the idea of partitioning variability within a dataset into "between-group" and "within-group" components, and then formally testing if the variability between groups is significantly different. Let's look at one-factor and two-factor ANOVA:

### One-Factor (or One-Way) ANOVA
. 1point3acres
#### Purpose:
To test if there are any statistically significant differences between the means of three or more independent (unrelated) groups.
. Waral dи,
#### Hypotheses:
- \(H_0\): The means of the different groups are equal.
- \(H_a\): At least one group mean is different from the others.

#### Procedure:
1. **Partition Variance**: Variance in the dataset is categorized into "between-group" and "within-group" variance.
2. **Calculate F-Statistic**: The F-statistic is calculated as the ratio of "between-group" variance to "within-group" variance.
3. **Compare to F-Distribution**: The calculated F-statistic is compared to the F-distribution with degrees of freedom determined by the sample sizes in the groups.

#### Key Output:
- F-statistic
- P-value

#### Decision:
- If the p-value is less than the significance level (\( \alpha \), often 0.05), reject the null hypothesis.. ----
. check 1point3acres for more.
#### Assumptions:
- Independence of observations
- Normal distribution within each group
- Homogeneity of variances (equal variances across groups)
.1point3acres
---

### Two-Factor (or Two-Way) ANOVA

#### Purpose:
To assess the individual and interactive effects of two categorical independent variables on an outcome variable.

#### Types:
1. **Two-Way ANOVA Without Replication**: One observation for each combination of factors.
2. **Two-Way ANOVA With Replication**: More than one observation for each combination of factors.

#### Hypotheses:
- \(H_0\): The means are equal across levels of each factor.-baidu 1point3acres
- \(H_a\): Means are different across levels of each factor.

There are three sets of hypothesis tests in a two-way ANOVA:
1. Effect of factor A on the outcome.
2. Effect of factor B on the outcome.
3. Interaction effect between factors A and B on the outcome.

#### Procedure:
1. **Partition Variance**: Variance is divided into variance due to factor A, variance due to factor B, interaction variance, and error variance.
2. **Calculate F-Statistics**: Separate F-statistics are calculated for the effects of Factor A, Factor B, and their interaction.
3. **Compare to F-Distribution**: Each F-statistic is compared to an F-distribution with appropriate degrees of freedom.

#### Key Output:
- Three F-statistics and their corresponding p-values (one for each source of variance: Factor A, Factor B, and interaction).. ----

#### Decision:
- Reject or fail to reject the null hypotheses based on p-values and the F-distribution.. check 1point3acres for more.

#### Assumptions:-baidu 1point3acres
- Independence of observations
- Normal distribution within each group
- Homogeneity of variances
. 1point3acres
Both types of ANOVA are powerful tools for understanding the factors that influence a given outcome variable. However, if the assumptions of ANOVA are not met, the results may not be reliable, and other types of statistical tests may be more appropriate.
回复

使用道具 举报

🔗
 楼主| 期待阳光 2023-9-2 22:21:23 | 只看该作者
全局:
SQL
In SQL, a Common Table Expression (CTE) is a named temporary result set that you can reference within a SELECT, INSERT, UPDATE, or DELETE statement. A CTE is defined using the WITH clause. It can be thought of as a derived table that exists just for the duration of the query. CTEs are particularly useful for breaking down complex queries into simpler parts, making your SQL code more readable and easier to maintain. Context:
WITH cte_name (column_name1, column_name2, ...). 1point3acres
AS (
  -- SQL query that generates the CTE data
)
-- SQL query that uses the CTE
回复

使用道具 举报

🔗
 楼主| 期待阳光 2023-12-19 10:13:04 来自APP | 只看该作者
全局:
Parameters vs  hyperparameters:
..
1. **Parameters:**
   - **Definition:** Parameters are the internal variables that the model learns from the training data. They are the coefficients in linear regression, the weights in a neural network, or the splits in a decision tree.. ----
   - **Role:** Parameters are optimized during the training process to minimize the difference between the model's predictions and the actual outcomes in the training data.
   - **Example:** In a linear regression model, the coefficients (weights) assigned to each feature are parameters.

2. **Hyperparameters:**
   - **Definition:** Hyperparameters are external configuration settings that are not learned from the data but are set prior to the training process. They control the overall behavior of the model.
   - **Role:** Hyperparameters are not learned from the data but are set by the data scientist or machine learning engineer. They influence the learning process and the structure of the model.
   - **Example:** Learning rate in gradient boosting, the number of layers in a neural network, or the depth of a decision tree are hyperparameters.

补充内容 (2023-12-19 11:40 +08:00):
.google  и
training set - train parameters
validation set - train hyper parameters
test set - test generalized error



回复

使用道具 举报

🔗
 楼主| 期待阳光 2023-12-19 11:39:43 | 只看该作者
全局:
How to use Confidence Interval for model comparison -- deep learning book Chapter 5 p31
Suppose we have two regression algorithms, Algorithm A and Algorithm B, and we want to compare their mean squared errors on a specific dataset.

Assume the MSE values for 100 data points are as follows:

- For Algorithm A: \(MSE_A = 20\) with a 95% confidence interval \([18, 22]\).
- For Algorithm B: \(MSE_B = 25\) with a 95% confidence interval \([23, 28]\).

Now, let's apply the comparison criterion:

"Algorithm A is better than Algorithm B if the upper bound of the 95% confidence interval for the MSE of Algorithm A is less than the lower bound of the 95% confidence interval for the MSE of Algorithm B.".

- The upper bound of the confidence interval for Algorithm A is 22.
- The lower bound of the confidence interval for Algorithm B is 23.

. 1point3acresSince \(22 < 23\), we can conclude, with 95% confidence, that Algorithm A is likely to have a lower mean squared error than Algorithm B. In the context of MSE, lower values are generally desirable because they indicate better model performance in terms of minimizing prediction errors.
回复

使用道具 举报

您需要登录后才可以回帖 登录 | 注册账号
隐私提醒:
  • ☑ 禁止发布广告,拉群,贴个人联系方式:找人请去🔗同学同事飞友,拉群请去🔗拉群结伴,广告请去🔗跳蚤市场,和 🔗租房广告|找室友
  • ☑ 论坛内容在发帖 30 分钟内可以编辑,过后则不能删帖。为防止被骚扰甚至人肉,不要公开留微信等联系方式,如有需求请以论坛私信方式发送。
  • ☑ 干货版块可免费使用 🔗超级匿名:面经(美国面经、中国面经、数科面经、PM面经),抖包袱(美国、中国)和录取汇报、定位选校版
  • ☑ 查阅全站 🔗各种匿名方法

本版积分规则

>
快速回复 返回顶部 返回列表