What is the purpose of testing a Machine Learning model?

  • To cluster the data
  • To evaluate the model's performance on unseen data
  • To label the data
  • To train the model
The purpose of testing a Machine Learning model is to evaluate its performance on unseen data. This helps in assessing how well the model generalizes to new instances.

The DBSCAN algorithm groups together points that are closely packed, forming clusters, and treats more isolated points as _________.

  • Centroids
  • Clusters
  • Noise
  • Outliers
In DBSCAN, more isolated points that don't belong to any cluster are treated as outliers.

How is the amount of variance explained related to Eigenvalues in PCA?

  • Eigenvalues are unrelated to variance
  • Eigenvalues represent the mean of the data
  • Larger eigenvalues explain more variance
  • Smaller eigenvalues explain more variance
In PCA, the amount of variance explained by each principal component is directly related to its corresponding eigenvalue. Larger eigenvalues mean that more variance is explained by that particular component.

The ___________ regression technique can be used when the relationship between the independent and dependent variables is not linear.

  • L1 Regularization
  • Logistic
  • Polynomial
  • Simple Linear
Polynomial Regression can model non-linear relationships between independent and dependent variables by transforming the predictors into a polynomial form, allowing for more complex fits.

How does adding regularization help in avoiding overfitting?

  • By adding noise to the training data
  • By fitting the model closely to the training data
  • By increasing model complexity
  • By reducing model complexity
Regularization helps in avoiding overfitting by "reducing model complexity." It adds a penalty to the loss function, constraining the weights and preventing the model from fitting too closely to the training data.

How can you tune hyperparameters in SVM to prevent overfitting?

  • Changing the color of hyperplane
  • Increasing data size
  • Reducing feature dimensions
  • Using appropriate kernel and regularization
Tuning hyperparameters like the choice of kernel and regularization helps in controlling model complexity to prevent overfitting in SVM.

Explain how a Decision Tree works in the context of Machine Learning.

  • Based on complexity, combines data at each node
  • Based on distance, groups data at each node
  • Based on entropy, splits data at each node
  • Based on gradient, organizes data at each node
A Decision Tree works by splitting the data into subsets based on feature values. This is done recursively at each node by selecting the feature that provides the best split according to a metric like entropy or Gini impurity. The process continues until specific criteria are met, creating a tree-like structure.

When it comes to classifying data points, the _________ algorithm considers the 'K' closest points to make a decision.

  • K-Nearest Neighbors (KNN)
  • Logistic Regression
  • Random Forest
  • Support Vector Machines
K-Nearest Neighbors (KNN) algorithm classifies a data point based on the majority class of its 'K' closest points in the dataset, using distance metrics to determine proximity.

The risk of overfitting can be increased if the same data is used for both _________ and _________ of the Machine Learning model.

  • evaluation, processing
  • training, testing
  • training, validation
  • validation, training
If the same data is used for both "training" and "testing," the model may perform well on that data but poorly on unseen data, leading to overfitting.

To detect multicollinearity in a dataset, one common method is to calculate the ___________ Inflation Factor (VIF).

  • Validation
  • Variable
  • Variance
  • Vector
The Variance Inflation Factor (VIF) is a measure used to detect multicollinearity. It quantifies how much a variable is inflating the standard errors due to its correlation with other variables. A high VIF indicates multicollinearity.