In which situations would you prefer a heatmap over a scatter plot?

  • When dealing with a single variable
  • When the data is non-numeric
  • When visualizing a time series
  • When visualizing the correlation between multiple variables
You would prefer a heatmap over a scatter plot when visualizing the correlation between multiple variables. A heatmap can visualize any number of variables at once and is particularly effective when the dataset contains many variables.

Which of the following types of analysis provides the least assumptions about data: EDA, CDA, or Predictive Modeling?

  • CDA
  • EDA
  • Predictive Modeling
  • They all make the same number of assumptions.
EDA makes the least assumptions about data. While CDA and Predictive Modeling typically require some assumptions about the data's distribution or the relationships between variables, EDA is a more open-ended exploration of the data's structure and patterns.

You have a scatter plot with a strong positive correlation, but a few points are far from the correlation line. What might these points represent?

  • Correlated data points
  • False positives
  • Normal data points
  • Outliers
In a scatter plot, points that are far away from the correlation line often represent outliers.

You create a histogram of a dataset and notice that the frequency count is very high on the far right of the distribution but drops significantly after that. What can be inferred from this?

  • Data has a negative skewness
  • Data has a positive skewness
  • Data is evenly distributed
  • Data is normally distributed
If the frequency count in a histogram is very high on the far right but drops significantly after that, it can indicate that the data has a positive skewness.

A data analyst needs to demonstrate the occurrence of outliers in a dataset using a plot. Which plot type would you recommend and why?

  • Bar graph
  • Box plot
  • Line graph
  • Scatter plot
The Box plot is ideal for demonstrating outliers in a dataset. The 'whiskers' in a box plot represent the range for the bulk of the data, and any data point that falls outside of this range is visually represented as an outlier.

The process of combining highly correlated variables into one is called _________.

  • Data Aggregation
  • Principal Component Analysis (PCA)
  • Standardization
  • Variance Inflation
When dealing with multicollinearity, one approach is to combine the correlated variables into one using a technique such as Principal Component Analysis (PCA). PCA creates new uncorrelated variables that capture the information of the original variables.

The ______ of a scatter plot may indicate the presence of outliers in the dataset.

  • correlation
  • scatter
  • slope
  • trend line
In a scatter plot, the scattering or spread of data points can help identify outliers. Points that are distant from the main concentration of data can indicate potential outliers.

In ____________, different models are used to estimate the missing values based on observed data.

  • Mean Imputation
  • Mode-based Imputation
  • Model-based Imputation
  • Multiple Imputation
This process is called model-based imputation. Different statistical models are used to estimate the missing values based on the observed (non-missing) data.

You have a dataset with missing values and you've chosen to use multiple imputation. However, the results after applying multiple imputation are not as expected. What factors might be causing this?

  • Both too few and too many imputations
  • The model used for imputation is perfect
  • Too few imputations
  • Too many imputations
If too few imputations are used in multiple imputation, the results may not be accurate. This may lead to an underestimation of standard errors and incorrect statistical inference. Increasing the number of imputations generally leads to more accurate results.

What type of data typically requires more complex statistical methods for analysis?

  • Categorical data
  • Continuous data
  • Discrete data
  • Ordinal data
Continuous data usually requires more complex statistical methods for analysis because it can take on any value within a certain range. This might require techniques like regression, hypothesis testing, and advanced graphical representations.