When data points are concentrated on the left and the tail is on the right, the distribution is said to be _______.

  • Negatively skewed
  • Normal
  • Positively skewed
  • Uniform
When data points are concentrated on the left and the tail is on the right, the distribution is said to be positively skewed or right-skewed. This is because the tail of the distribution points towards the positive end of the axis.

Why is residual analysis important in regression models?

  • To check the assumptions of the regression model
  • To determine the slope of the regression line
  • To estimate the parameters of the model
  • To predict the dependent variable
Residual analysis is important because it helps us to validate the assumptions of the regression model, such as linearity, independence, normality, and equal variance (homoscedasticity). This is crucial for the reliability and validity of the regression model.

What is the significance of the total probability rule?

  • It is a rule for determining the probability of dependent events
  • It is used to calculate conditional probabilities
  • It is used to calculate the probability of mutually exclusive events
  • It provides a way to break down probabilities of complex events into simpler ones
The Total Probability Rule provides a way to compute the probability of an event from the probabilities of that event occurring within disjoint subsets of the sample space. It essentially allows you to break down the probability of complex events into simpler or more basic component events.

In a Chi-square test for goodness of fit, the degrees of freedom are calculated as the number of categories minus ________.

  • one
  • the number of samples
  • three
  • two
In a Chi-square test for goodness of fit, the degrees of freedom are calculated as the number of categories minus one. This reflects the number of values in the final calculation that are free to vary.

How does bin size affect a histogram representation?

  • Bin size changes the shape of the histogram
  • Bin size does not affect the histogram
  • Larger bins make the histogram more detailed
  • Smaller bins make the histogram more detailed
The choice of bin size in a histogram can greatly affect the resulting visualization. If the bins are too large, important features of the data may be obscured. If the bins are too small, the histogram may appear too 'noisy' and it may be difficult to interpret underlying patterns. Thus, the choice of bin size can indeed change the perceived shape of the histogram.

What is multicollinearity and how does it affect simple linear regression?

  • It is the correlation between dependent variables and it has no effect on regression
  • It is the correlation between errors and it makes the regression model more accurate
  • It is the correlation between independent variables and it can cause instability in the regression coefficients
  • It is the correlation between residuals and it causes bias in the regression coefficients
Multicollinearity refers to a high correlation among independent variables in a regression model. It does not reduce the predictive power or reliability of the model as a whole, but it can cause instability in the estimation of individual regression coefficients, making them difficult to interpret.

The distribution of all possible sample means is known as a __________.

  • Normal Distribution
  • Population Distribution
  • Sampling Distribution
  • Uniform Distribution
The sampling distribution in statistics is the probability distribution of a given statistic based on a random sample. For a statistic that is calculated from a sample, each different sample could (and likely will) provide a different value of that statistic. The sampling distribution shows us how those calculated statistics would be distributed.

How is 'K-means' clustering different from 'hierarchical' clustering?

  • Hierarchical clustering creates a hierarchy of clusters, while K-means does not
  • Hierarchical clustering uses centroids, while K-means does not
  • K-means requires the number of clusters to be defined beforehand, while hierarchical clustering does not
  • K-means uses a distance metric to group instances, while hierarchical clustering does not
K-means clustering requires the number of clusters to be defined beforehand, while hierarchical clustering does not. Hierarchical clustering forms a dendrogram from which the user can choose the number of clusters based on the problem requirements.

Under what conditions does a binomial distribution approximate a normal distribution?

  • When the events are not independent
  • When the number of trials is large and the probability of success is not too close to 0 or 1
  • When the number of trials is small
  • When the probability of success changes with each trial
The binomial distribution approaches the normal distribution as the number of trials gets large, provided that the probability of success is not too close to 0 or 1. This is known as the De Moivre–Laplace theorem.

What does the F-statistic signify in an ANOVA test?

  • The ratio of between-group variability to within-group variability
  • The ratio of total variability to within-group variability
  • The ratio of within-group variability to between-group variability
  • The ratio of within-group variability to total variability
In an ANOVA test, the F-statistic is the ratio of the between-group variability to the within-group variability. In other words, it measures how much the means of each group vary between the groups, compared to how much they vary within each group. A larger F-statistic implies a greater degree of difference between the group means.