The process of training a Machine Learning model involves using a dataset known as the _________ set, while evaluating it involves the _________ set.

  • testing, validation
  • training, testing
  • validation, testing
  • validation, training
In supervised learning, a "training" set is used to train the model, and a "testing" set is used to evaluate its predictive performance on unseen data.

How can centering variables help in interpreting interaction effects in Multiple Linear Regression?

  • By increasing model accuracy
  • By increasing prediction speed
  • By reducing multicollinearity between main effects and interaction terms
  • By simplifying the model
Centering variables (subtracting the mean) can reduce multicollinearity between main effects and interaction terms, making it easier to interpret the individual and combined effects of the variables.

In a case where your regression model is suffering from high variance, what regularization technique might you apply, and why?

  • Increase model complexity
  • L1 regularization
  • L2 regularization (Ridge)
  • Reduce model complexity
High variance in a regression model often signals overfitting, where the model performs well on training data but poorly on unseen data. L2 regularization (Ridge regression) can help by penalizing large coefficients, reducing overfitting, and improving generalization.

How does the objective function differ between Ridge, Lasso, and ElasticNet?

  • No difference
  • Ridge and Lasso have the same objective
  • Ridge uses L1, Lasso uses L2, ElasticNet uses neither
  • Ridge uses L2, Lasso uses L1, ElasticNet uses both
Ridge's objective function includes an L2 penalty, Lasso's includes an L1 penalty, and ElasticNet's includes both L1 and L2 penalties.

How do Precision and Recall trade-off in a classification problem, and when might you prioritize one over the other?

  • Increasing Precision decreases Recall, prioritize Precision when false positives are costly
  • Increasing Precision increases Recall, prioritize Recall when false positives are costly
  • Precision and Recall are independent, no trade-off
  • nan
Precision and Recall often trade-off; increasing one can decrease the other. You might prioritize Precision when false positives are more costly (e.g., spam detection) and Recall when false negatives are more costly (e.g., fraud detection).

While performing Cross-Validation, you notice a significant discrepancy between training and validation performance in each fold. What might be the reason, and how would you address it?

  • All of the above
  • Data leakage; ensure proper separation between training and validation
  • Overfitting; reduce model complexity
  • Underfitting; increase model complexity
A significant discrepancy between training and validation performance could result from overfitting, underfitting, or data leakage. Addressing it requires identifying the underlying issue and taking appropriate action, such as reducing/increasing model complexity for overfitting/underfitting or ensuring proper separation between training and validation to prevent leakage.

What is the primary purpose of using ensemble methods in machine learning?

  • To combine multiple weak models to form a strong model
  • To focus on a single algorithm
  • To reduce computational complexity
  • To use only the best model
Ensemble methods combine the predictions from multiple weak models to form a more robust and accurate model. By leveraging the strength of multiple models, they typically achieve better generalization and performance than using a single model.

Cross-validation, such as _______-fold cross-validation, can help in detecting and preventing overfitting.

  • 10
  • 3
  • 5
  • any number
Any number of folds can be used in cross-validation, although commonly used numbers include 5 and 10. Cross-validation helps in model validation and prevents overfitting.

Can you explain the main concept behind boosting algorithms?

  • Boosting always uses Random Forest
  • Boosting combines models sequentially, giving more weight to misclassified instances
  • Boosting focuses on the strongest predictions
  • Boosting involves reducing model complexity
Boosting is an ensemble method where models are combined sequentially, with each model focusing more on the instances that were misclassified by the previous models. This iterative process helps in correcting the mistakes of earlier models, leading to improved performance.

The point in the ROC Curve where the True Positive Rate equals the False Positive Rate is known as the __________ point.

  • Break-even
  • Equilibrium
  • Random
  • nan
The Break-even point on the ROC Curve is where the True Positive Rate equals the False Positive Rate. This point represents a balance between sensitivity and specificity.