What is the primary purpose of factor analysis in data science?

  • To categorize data
  • To classify data
  • To identify underlying variables (factors)
  • To predict future outcomes
Factor analysis is a statistical method used to describe variability among observed, correlated variables in terms of a potentially lower number of unobserved variables called factors. Its primary purpose is to identify the underlying structure and relationships within a set of variables.

What does it mean when a confidence interval includes the value zero?

  • The population mean is likely to be zero
  • The sample mean is zero
  • There is no effect in the population
  • nan
If a confidence interval for a mean difference or an effect size includes zero, it suggests that there is no effect in the population and that the observed effect in the sample is likely due to sampling error.

Can you provide a practical example of where the Law of Large Numbers is applied?

  • Insurance companies use the Law of Large Numbers to predict claim amounts.
  • It's used to calculate the speed of light.
  • The Law of Large Numbers is only theoretical and has no practical applications.
  • The Law of Large Numbers is used to predict lottery numbers.
The Law of Large Numbers has many practical applications. For example, insurance companies use it to predict future claim amounts. The law allows them to predict losses and to set premiums in a way that ensures profitability, by basing predictions on large aggregations of independent or nearly independent losses.

What effect does a high leverage point have on a multiple linear regression model?

  • It can significantly affect the estimate of the regression coefficients
  • It does not affect the model
  • It increases the R-squared value
  • It leads to homoscedasticity
High leverage points are observations with extreme values on the predictor variables. They can have a disproportionate influence on the estimation of the regression coefficients, potentially leading to a less reliable model.

How does multicollinearity affect the interpretation of regression coefficients?

  • It has no effect on the interpretation of the coefficients.
  • It increases the value of the coefficients.
  • It makes the coefficients less interpretable and reliable.
  • It makes the coefficients more interpretable and reliable.
Multicollinearity can cause large changes in the estimated regression coefficients for small changes in the data. Hence, it makes the coefficients less reliable and interpretable.

If two events are independent, what is the conditional probability of one given the other?

  • 0
  • 1
  • Equal to the probability of the given event
  • Undefined
If two events are independent, the conditional probability of one event given the other is simply the probability of the event itself. This is because in independent events, the occurrence of one event does not affect the occurrence of the other event.

In what situation could a "Type II" error occur during hypothesis testing?

  • When the alternative hypothesis is false
  • When the null hypothesis is false but not rejected
  • When the null hypothesis is rejected
  • When the null hypothesis is true
A Type II error, also known as a false negative, occurs when the null hypothesis is false, but we fail to reject it.

Under what circumstances can the mode of a data set be irrelevant or misleading?

  • When the data is continuous
  • When the data set is large
  • When the data set is small
  • When there are multiple modes
The mode can be misleading or irrelevant especially with continuous data. Since the mode is the most frequently occurring value, with continuous data the frequency of each value is often the same (i.e., 1), hence it becomes difficult to define a mode in a traditional sense.

A Variance Inflation Factor (VIF) greater than 5 indicates a high degree of _______ among the predictors.

  • correlation
  • distribution
  • multicollinearity
  • variance
A VIF greater than 5 is often taken as an indication of high multicollinearity among the predictors in a regression model. This could lead to imprecise and unreliable estimates of the regression coefficients.

How does the 'elbow method' help in determining the optimal number of clusters in K-means clustering?

  • By calculating the average distance between all pairs of clusters
  • By comparing the silhouette scores for different numbers of clusters
  • By creating a dendrogram of clusters
  • By finding the point in the plot of within-cluster sum of squares where the decrease rate sharply shifts
The elbow method involves plotting the explained variation as a function of the number of clusters and picking the elbow of the curve as the number of clusters to use. This 'elbow' is the point representing the optimal number of clusters at which the within-cluster sum of squares (WCSS) doesn't decrease significantly with each iteration.