What is the relationship between the mean and the standard deviation in a normal distribution?

  • The mean is always larger than the standard deviation
  • The mean is the midpoint of the distribution, and the standard deviation measures the spread
  • The standard deviation is always larger than the mean
  • There is no relationship between the mean and the standard deviation
In a normal distribution, the mean is the center of the distribution and represents the "average" value. The standard deviation measures the dispersion around the mean. Roughly 68% of the data falls within one standard deviation of the mean in a normal distribution.

The conditional probability of A given B is denoted as ________.

  • P(A + B)
  • P(A / B)
  • P(A B)
  • P(A ∩ B)
The conditional probability of A given B is denoted as P(A

What is the primary goal of random sampling?

  • To always select the same individuals
  • To ensure that every member of the population has an equal chance of being selected
  • To select individuals who are likely to give the desired results
  • To select the individuals who are easiest to reach
The primary goal of random sampling is to ensure that every member of the population has an equal chance of being selected. This helps to reduce bias and increase the likelihood that the sample is representative of the population, which makes the results more valid and generalizable.

_______ regression is a method used to handle multicollinearity by adding a degree of bias to the regression estimates.

  • Logistic
  • Polynomial
  • Ridge
  • Simple linear
Ridge regression handles multicollinearity by introducing a degree of bias to the regression estimates, reducing their variance, and making them more reliable.

What's the difference between a histogram and a bar plot?

  • Bar plots are for continuous data, histograms for categorical data
  • Both are for continuous data only
  • Histograms are for continuous data, bar plots for categorical data
  • There is no difference
The main difference between a histogram and a bar plot is the type of data they represent. A histogram is used for continuous data, where the bins represent ranges of data, while a bar plot is used for categorical data to compare the frequency or count of different categories.

What is the error term in a simple linear regression model?

  • It is the dependent variable
  • It is the difference between the observed and predicted values
  • It is the independent variable
  • It is the slope of the regression line
The error term in a simple linear regression model is the difference between the observed and predicted values. It captures the variability in the dependent variable that is not explained by the independent variable in the model.

What can be inferred if the residuals are not randomly distributed in the residual plot?

  • The data has no outliers
  • The data is perfectly linear
  • The linear regression model is a perfect fit for the data
  • The linear regression model is not a good fit for the data
If the residuals are not randomly distributed (e.g., if they form a pattern), it suggests that the linear regression model is not a good fit for the data. This could be because the relationship between the variables is not linear, or because the data exhibits heteroscedasticity (unequal variances of errors), among other reasons.

What type of data is used in the Chi-square test for goodness of fit?

  • Categorical data
  • Continuous data
  • Interval data
  • Ordinal data
The Chi-square test for goodness of fit is used with categorical data. It compares the observed frequencies in each category with the frequencies we would expect to see if the data followed the theoretical distribution.

What is the null hypothesis in the Mann-Whitney U test?

  • The groups have different variances
  • The groups have equal variances
  • There is a significant difference between the groups
  • There is no significant difference between the groups
In the Mann-Whitney U test, the null hypothesis is that there is no significant difference between the groups. More specifically, it states that the probability that a randomly selected value from the first group is greater than a randomly selected value from the second group is equal to 0.5.

How does sample size affect the width of a confidence interval?

  • Increasing the sample size decreases the width of the confidence interval
  • Increasing the sample size has no effect on the width of the confidence interval
  • Increasing the sample size increases the width of the confidence interval
  • The relationship between sample size and the width of the confidence interval is unpredictable
Increasing the sample size decreases the width of the confidence interval. The larger the sample size, the more information you have, and thus the less uncertainty (which translates into a smaller standard error and narrower confidence interval).