EDA generally precedes ________ in the data analysis process.
- Confirmatory Data Analysis
- Data cleaning
- Data collection
- Predictive Modeling
EDA generally precedes Confirmatory Data Analysis (CDA) in the data analysis process. While EDA is all about exploring data to find patterns and relationships, CDA is about confirming or falsifying existing hypotheses.
When applying multiple imputation, increasing the number of imputations can help reduce the ____________.
- Mean
- Mode
- Sampling error
- Standard deviation
When applying multiple imputation, increasing the number of imputations can help reduce the sampling error. More imputations allow for a better representation of the uncertainty due to missingness, resulting in more accurate standard errors and confidence intervals.
Can feature selection improve the computational efficiency of a machine learning model?
- Depends on the dataset
- Depends on the model
- No
- Yes
Yes, feature selection can improve the computational efficiency of a machine learning model by reducing the number of features it needs to process.
You are analyzing a dataset where some missing values have been replaced using mean imputation. What effect might this have on the variance of the data?
- It could cause overfitting
- It could create multicollinearity
- It could decrease the variance
- It could increase the variance
When missing values are replaced using mean imputation, it could decrease the variance of the data. This is because imputed values are just the mean of observed values and do not add any variability. Therefore, the overall variability of the data could be underestimated, leading to biased estimates.
You're working with a dataset where two features, 'age' and 'years of experience', have a high correlation. Which problem does this situation exemplify?
- Data leakage
- Multicollinearity
- Overfitting
- Underfitting
This situation exemplifies multicollinearity, a condition where two or more predictors in a multiple regression model are highly correlated. This high correlation means that 'age' and 'years of experience' provide similar information in predicting the dependent variable.
What type of data can only take on discrete values?
- Categorical data
- Continuous data
- Discrete data
- Ordinal data
Discrete data can only take on distinct, separate values. It can't be made more precise by further measurement or counting. For example, the number of students in a class would be discrete data.
You are given a dataset with a high number of features. The computational resources are limited. What feature selection method might you consider?
- Backward elimination
- Filter methods
- Forward selection
- Wrapper methods
Given limited computational resources, filter methods might be a good choice. These methods are less computationally expensive than wrapper methods as they do not involve the use of any specific machine learning algorithm. Instead, they rank features based on statistical measures and remove irrelevant features based on a certain threshold or number of top features to keep.
_____ are used to indicate different values in a heatmap.
- Colors
- Lines
- Shapes
- Sizes
Colors are used to indicate different values in a heatmap. The color scale represents the magnitude of the variable, with different color gradients representing different value ranges.
How do outliers affect the standard deviation of a dataset?
- Can either increase or decrease the standard deviation
- Decrease the standard deviation
- Does not affect the standard deviation
- Increase the standard deviation
Outliers can significantly increase the standard deviation, as the standard deviation is sensitive to extreme values. This is because the standard deviation squares the differences from the mean, making it more reactive to values far from the mean.
How do outliers affect the performance of machine learning models?
- Decrease model accuracy
- Increase model accuracy
- Increase model precision
- Increase model recall
Outliers can significantly affect the performance of machine learning models, often leading to decreased accuracy. This is because they can cause the model to learn based on these anomalies rather than the underlying data pattern.
How can extreme outliers impact the interpretation of the skewness of a dataset?
- Can either increase or decrease the skewness
- Decrease the skewness
- Does not affect the skewness
- Increase the skewness
The skewness of a distribution is a measure of the extent and direction of asymmetry. Extreme outliers can either increase or decrease skewness depending on which tail they lie in. If the outliers are greater than the mean, skewness will be increased. If less, skewness will be decreased.
In a Normal Distribution, approximately 95% of the data falls within _____ standard deviations of the mean.
- 1
- 2
- 3
- 4
In a Normal Distribution, approximately 95% of the data falls within 2 standard deviations of the mean.