In EDA, which method can help in understanding how a single variable is distributed across various categories or groups?

  • Histogram
  • Box Plot
  • Scatter Plot
  • Bar Plot
A bar plot is used to visualize the distribution of a single variable across different categories or groups. It displays the data in rectangular bars, making it easy to compare and understand how the variable is distributed among the categories. Commonly used in Exploratory Data Analysis (EDA).

You're working with a dataset containing sales data from various regions. You want to identify sales patterns, seasonal trends, and anomalies. Which EDA techniques and visualization tools would be best suited for this?

  • Scatter plots and t-SNE
  • Box plots and bar charts
  • Time series plots and heatmaps
  • Histograms and parallel coordinates
For exploring sales patterns and seasonal trends, time series plots and heatmaps are excellent choices. Time series plots can reveal trends over time, and heatmaps can show correlations between different regions and sales data, helping identify anomalies and patterns.

Which method in transfer learning involves freezing the earlier layers of a pre-trained model and only training the latter layers for the new task?

  • Fine-tuning
  • Knowledge Transfer
  • Feature Extraction
  • Weight Sharing
The method in transfer learning that involves freezing the earlier layers of a pre-trained model and only training the latter layers for the new task is known as fine-tuning. Fine-tuning allows the model to retain the knowledge from the source task while adapting its later layers for the specific requirements of the target task. This approach is common in transfer learning scenarios.

While working with a dataset about car sales, you discover that the "Brand" column has many brands with very low frequency. To avoid having too many sparse categories, which technique can you apply to the "Brand" column?

  • One-Hot Encoding
  • Label Encoding
  • Brand grouping based on frequency
  • Principal Component Analysis (PCA)
To handle low-frequency categories in the "Brand" column, you can group the brands based on their frequency. This reduces the number of sparse categories and can improve model performance. You can also consider techniques like label encoding or one-hot encoding, but they might not be ideal for low-frequency categories. PCA is used for dimensionality reduction and not for handling categorical variables.

A neural network without any hidden layers is typically referred to as a _______.

  • Deep Neural Network
  • Shallow Neural Network
  • Multilayer
  • Perceptron
A neural network without any hidden layers is often referred to as a "Perceptron." It consists of only the input and output layers, and it's the simplest form of a neural network.

To avoid data leakage during transformation, one should fit the scaler on the _______ set and transform both the training and test sets.

  • Training
  • Validation
  • Test
  • Entire Dataset
To prevent data leakage, it's essential to fit a scaler on the training set (Option A) and then apply the same transformation to both the training and test sets. This ensures that the test set remains independent of the training data.

Before deploying a model into production in the Data Science Life Cycle, it's essential to have a _______ phase to test the model's real-world performance.

  • Training phase
  • Deployment phase
  • Testing phase
  • Validation phase
Before deploying a model into production, it's crucial to have a testing phase to evaluate the model's real-world performance. This phase assesses how the model performs on unseen data to ensure its reliability and effectiveness.

You're analyzing a dataset with the heights of individuals. While the mean height is 165 cm, you notice a few heights recorded as 500 cm. These values are likely:

  • Data entry errors
  • Outliers
  • Missing data
  • Measurement errors
The heights recorded as 500 cm are likely outliers in the dataset. Outliers are data points that significantly differ from the majority of the data and may indicate measurement errors or anomalies. It's important to identify and handle outliers appropriately during data analysis.

In time series forecasting, which method captures both trend and seasonality in the data?

  • Moving Average
  • Exponential Smoothing
  • ARIMA (AutoRegressive Integrated Moving Average)
  • Exponential Moving Average
ARIMA (AutoRegressive Integrated Moving Average) captures both trend and seasonality in time series data. It combines autoregressive, differencing, and moving average components to model complex time series patterns, making it a powerful method for forecasting data with seasonal and trend components.

Which Python library is specifically designed for statistical data visualization and is built on top of Matplotlib?

  • Seaborn
  • Pandas
  • Numpy
  • Scikit-learn
Seaborn is a Python library built on top of Matplotlib, designed for statistical data visualization. It provides a high-level interface for creating informative and attractive statistical graphics, making it a valuable tool for data analysis and visualization.