As a data scientist, you've realized that your dataset contains missing values. How would you handle this situation as part of your EDA process?

  • Always replace missing values with the mean or median
  • Choose an appropriate imputation method depending on the nature of the data and the type of missingness
  • Ignore the missing values and proceed with analysis
  • Remove all instances with missing values
Handling missing values is an important part of the EDA process. The method used to handle them depends on the nature of the data and the type of missingness (MCAR, MAR, or NMAR). Various imputation methods can be used, such as mean/median/mode imputation for MCAR or MAR data, and advanced imputation methods like regression imputation, multiple imputation, or model-based methods for NMAR data.

If the variance of a data set is zero, then all data points are ________.

  • Equal
  • Infinite
  • Negative
  • Positive
If the "Variance" of a data set is zero, then all data points are "Equal". Variance is a measure of how far a set of numbers is spread out from their average value. A variance of zero indicates that all the values within a set of data are identical.

A market research survey collects data on customer age, gender, and preference for a product (Yes/No). Identify the types of data present in this survey.

  • Age: continuous, Gender: nominal, Preference: ordinal
  • Age: nominal, Gender: ordinal, Preference: interval
  • Age: ordinal, Gender: interval, Preference: ratio
  • Age: ratio, Gender: ordinal, Preference: nominal
Age is a continuous data type because it can take on any value within a range. Gender is nominal as it's categorical with no order or priority. Preference is ordinal as it's categorical with a clear order (Yes is preferred to No).

You notice that using the Z-score method for a particular data set is yielding too many outliers. What modifications can you make to the method to reduce the number of outliers detected?

  • Decrease the Z-score threshold
  • Increase the Z-score threshold
  • Use the IQR method instead
  • Use the modified Z-score method instead
Increasing the Z-score threshold will mean fewer points will exceed it, thus fewer outliers will be identified.

How can EDA assist in identifying errors or anomalies in the dataset?

  • By conducting a statistical test of normality
  • By creating a correlation matrix of the variables
  • By running the dataset through a predefined ML model
  • By summarizing and visualizing the data, which can reveal unexpected values or patterns
EDA, especially through summarizing and visualizing data, can assist in identifying errors or anomalies in the dataset. Graphical representations of data often make it easier to spot unexpected values, patterns, or aberrations that may not be apparent in the raw data.

You are given a dataset for an upcoming data analysis project. What initial EDA steps would you take before moving to model building?

  • Explore the structure of the dataset, summarize the data, and create visualizations
  • Perform a detailed statistical analysis
  • Run a quick ML model to test the data
  • Start cleaning and wrangling the data
Before moving to model building, it's important to first understand the dataset you're working with. The initial EDA steps would typically include exploring the structure of the dataset, summarizing the data (such as calculating central tendency measures and dispersion), and creating visualizations to uncover patterns, trends, and relationships.

How does standardization (z-score) affect the distribution of data?

  • It doesn't affect the shape of the distribution
  • It makes the distribution normal
  • It makes the distribution uniform
  • It skews the distribution
Standardization does not change the shape of the distribution of the feature; rather, it standardizes the scale. This means that it doesn't change the distribution's skewness or kurtosis but it does center the data around zero with a standard deviation of 1.

You are analyzing the number of calls received by a call center per hour. Which distribution would be most suitable for modeling this data and why?

  • Binomial Distribution because it represents the number of successes in a given number of trials
  • Normal Distribution because it represents continuous data
  • Poisson Distribution because it models the number of events occurring in a fixed interval of time
  • Uniform Distribution because all outcomes are equally likely
The Poisson Distribution is most suitable for modeling the number of calls received by a call center per hour because it models the number of events (calls) occurring in a fixed interval of time (per hour).

Consider a data distribution with a positive skewness and a high kurtosis. What does this scenario indicate about the distribution?

  • It has a symmetrical distribution.
  • It has evenly spread out values.
  • It has many values clustered around the left tail with potential outliers.
  • It has many values clustered around the right tail with potential outliers.
Positive skewness and high kurtosis imply that the data is heavily tailed to the right and the peak is sharp. Most of the data values are concentrated around the left tail, but there are potential outliers towards the more positive values.

What range of values does a dataset typically have after Min-Max scaling?

  • -1 to 1
  • 0 to 1
  • Depends on the dataset
  • Depends on the feature
Min-Max scaling transforms features by scaling each feature to a given range. The default range for the Min-Max scaling technique is 0 to 1. Therefore, after Min-Max scaling, the dataset will typically have values ranging from 0 to 1.