Given a machine learning algorithm that is highly sensitive to the range of input values, which scaling technique should you implement?

  • Min-Max scaling because it scales all values between 0 and 1
  • No scaling, as the original data values should be maintained
  • Robust scaling because it is not affected by outliers
  • Z-score standardization because it creates a normal distribution
Min-Max scaling is suitable when the algorithm is sensitive to the range of input values, as it scales all feature values into a specified range (usually 0-1). This ensures that all features have the same scale.

_____ data provides numerical measurements and it can be broken down into two subcategories: continuous and discrete.

  • Nominal
  • Ordinal
  • Qualitative
  • Quantitative
Quantitative data provides numerical measurements and it can be divided into two types: continuous (data that can take any value within a range) and discrete (data that can only take certain values).

In what scenario would a Poisson Distribution be a better fit than a Normal Distribution?

  • When modeling the number of times an event occurs in a fixed interval
  • When the data are continuous
  • When the data are negatively skewed
  • When the data are positively skewed
A Poisson Distribution would be a better fit when modeling the number of times an event occurs in a fixed interval of time or space. The Poisson Distribution is discrete while the Normal Distribution is continuous.

What is the Interquartile Range (IQR)?

  • The average spread of the data
  • The range of all the data
  • The range of the middle 50% of the data
  • The spread of the most common data
The Interquartile Range (IQR) is the "Range of the middle 50% of the data". It is calculated as the difference between the upper quartile (Q3) and the lower quartile (Q1).

Continuous data typically can be divided into which two main types?

  • Discrete and ordinal data
  • Interval and ratio data
  • Ordinal and nominal data
  • Qualitative and quantitative data
Continuous data can typically be divided into two main types: interval and ratio data. Interval data have a consistent scale but no true zero, while ratio data have a consistent scale and a true zero.

The square root of the ________ gives the standard deviation of a data set.

  • Mean
  • Median
  • Range
  • Variance
The "Variance" of a dataset is the average of the squared differences from the mean. The "Standard Deviation" is the square root of the variance. This means it's in the same unit as the data, which helps us understand the dispersion better.

Mishandling missing data can lead to a high level of ________, impacting model performance.

  • bias
  • precision
  • recall
  • variance
If missing data is handled improperly, it can lead to biased training data, which can cause the model to learn incorrect or irrelevant patterns and, as a result, adversely affect its performance.

How does multiple imputation handle missing data?

  • It deletes rows with missing data
  • It estimates multiple values for each missing value
  • It fills missing data with mode values
  • It replaces missing data with a single value
Multiple imputation estimates multiple values for each missing value, instead of filling in a single value for each missing point. It reflects the uncertainty around the true value and provides more realistic estimates.

Your EDA reveals a non-normal distribution of data in your dataset. How might this insight affect your choice of machine learning models or algorithms?

  • You should always normalize your data
  • You should use only non-parametric models
  • You should use only unsupervised learning models
  • Your choice of ML models might be influenced, as some models make certain assumptions about the data distribution
The distribution of data can influence the choice of machine learning models or algorithms. Some models, such as linear and logistic regression, make certain assumptions about the data distribution (i.e., they expect the input or output to be normally distributed). If these assumptions are violated, the model may perform poorly. Therefore, understanding the data distribution can guide you in choosing the most appropriate models or in deciding whether to transform your data.

Which type of correlation is based on ranks and perfect for ordinal data?

  • Kendall's Tau
  • Pearson's correlation
  • Point-Biserial Correlation
  • Spearman's correlation
Spearman's correlation, also known as Spearman's rank correlation, is based on ranks and is perfect for ordinal data. It assesses how well the relationship between two variables can be described using a monotonic function. It is less sensitive to outliers and non-linear relationships compared to Pearson's correlation.