How does Robust scaling minimize the effect of outliers?
- By ignoring them during the scaling process
- By removing the outliers
- By scaling based on the median and interquartile range instead of mean and variance
- By transforming the outliers
Robust scaling minimizes the effects of outliers by using the median and the interquartile range for scaling, instead of the mean and variance used by standardization. The interquartile range is the range between the 1st quartile (25th percentile) and the 3rd quartile (75th percentile). As the median and interquartile range are not affected by outliers, this method is robust to them.
Which measure of dispersion is defined as the difference between the largest and smallest values in a data set?
- Interquartile Range (IQR)
- Range
- Standard Deviation
- Variance
The "Range" is the measure of dispersion that is defined as the difference between the largest and smallest values in a data set.
The missing data mechanism where missingness is related only to the observed data is referred to as _________.
- All missing data
- MAR
- MCAR
- NMAR
In MAR (Missing at Random), the missingness is related only to the observed data.
You are given a dataset for an upcoming data analysis project. What initial EDA steps would you take before moving to model building?
- Explore the structure of the dataset, summarize the data, and create visualizations
- Perform a detailed statistical analysis
- Run a quick ML model to test the data
- Start cleaning and wrangling the data
Before moving to model building, it's important to first understand the dataset you're working with. The initial EDA steps would typically include exploring the structure of the dataset, summarizing the data (such as calculating central tendency measures and dispersion), and creating visualizations to uncover patterns, trends, and relationships.
How does standardization (z-score) affect the distribution of data?
- It doesn't affect the shape of the distribution
- It makes the distribution normal
- It makes the distribution uniform
- It skews the distribution
Standardization does not change the shape of the distribution of the feature; rather, it standardizes the scale. This means that it doesn't change the distribution's skewness or kurtosis but it does center the data around zero with a standard deviation of 1.
You are analyzing the number of calls received by a call center per hour. Which distribution would be most suitable for modeling this data and why?
- Binomial Distribution because it represents the number of successes in a given number of trials
- Normal Distribution because it represents continuous data
- Poisson Distribution because it models the number of events occurring in a fixed interval of time
- Uniform Distribution because all outcomes are equally likely
The Poisson Distribution is most suitable for modeling the number of calls received by a call center per hour because it models the number of events (calls) occurring in a fixed interval of time (per hour).
Consider a data distribution with a positive skewness and a high kurtosis. What does this scenario indicate about the distribution?
- It has a symmetrical distribution.
- It has evenly spread out values.
- It has many values clustered around the left tail with potential outliers.
- It has many values clustered around the right tail with potential outliers.
Positive skewness and high kurtosis imply that the data is heavily tailed to the right and the peak is sharp. Most of the data values are concentrated around the left tail, but there are potential outliers towards the more positive values.
What is the primary goal of Exploratory Data Analysis (EDA)?
- To confirm a pre-existing hypothesis
- To create an aesthetic representation of the data
- To make precise predictions about future events
- To understand the underlying structure of the data
The primary goal of EDA is to understand the underlying structure of the data, including distribution, variability, and relationships among variables. EDA allows analysts to make informed decisions about further data processing steps and analysis.
Data that follows a _____ Distribution has its values spread evenly across the range of possible outcomes.
- Binomial
- Normal
- Poisson
- Uniform
Data that follows a Uniform Distribution has its values spread evenly across the range of possible outcomes.
What is a key difference between qualitative data and quantitative data when it comes to analysis methods?
- All types of data are analyzed in the same way
- Qualitative data is always easier to analyze
- Qualitative data typically requires textual analysis, while quantitative data can be analyzed mathematically
- Quantitative data can't be used for statistical analysis
Qualitative data often requires textual or thematic analysis, categorizing the data based on traits or characteristics. Quantitative data, being numerical, can be analyzed using mathematical or statistical methods.