What is overfitting, and why is it a problem in Machine Learning models?

  • Fitting a model too loosely to training data
  • Fitting a model too well to training data, ignoring generalization
  • Ignoring irrelevant features
  • Including too many variables
Overfitting occurs when a model fits the training data too well, capturing noise rather than the underlying pattern. This leads to poor generalization to new data, resulting in suboptimal predictions on unseen data.

Describe the relationship between the Logit function, Odds Ratio, and the likelihood function in Logistic Regression.

  • The Logit function is used for multi-class, Odds Ratio for binary, likelihood for regression
  • The Logit function maps probabilities to log-odds, Odds Ratio quantifies effect on odds, likelihood function is used for estimation
  • The Logit function maps probabilities to odds, Odds Ratio quantifies effect on odds, likelihood function maximizes probabilities
  • They are unrelated
In Logistic Regression, the Logit function maps probabilities to log-odds, the Odds Ratio quantifies the effect of predictors on odds, and the likelihood function is used to estimate the model parameters by maximizing the likelihood of observing the given data.

Explain how Ridge and Lasso handle multicollinearity among the features.

  • Both eliminate correlated features
  • Both keep correlated features
  • Ridge eliminates correlated features; Lasso keeps them
  • Ridge keeps correlated features; Lasso eliminates them
Ridge regularization keeps correlated features but shrinks coefficients; Lasso can eliminate some by setting coefficients to zero.

What are some common applications for each of the four types of Machine Learning: Supervised, Unsupervised, Semi-Supervised, and Reinforcement?

  • Specific to finance
  • Specific to healthcare
  • Specific to manufacturing
  • Varies based on the problem domain
The applications for these types of Machine Learning vary and can be tailored to various problem domains, not confined to specific industries.

What is the difference between simple linear regression and multiple linear regression?

  • Number of dependent variables
  • Number of equations
  • Number of independent variables
  • Number of observations
Simple linear regression involves one independent variable to predict a dependent variable, whereas multiple linear regression uses two or more independent variables for prediction. The inclusion of more variables in multiple linear regression allows for more complex models and can lead to a better understanding of the relationships between variables.

The performance of an LDA model can be evaluated using ___________, which considers both within-class and between-class variances.

  • accuracy metrics
  • error rate
  • feature selection
  • principal components
"Accuracy metrics" that consider both within-class and between-class variances can be used to evaluate the performance of an LDA model. It gives a comprehensive view of how well the model has separated the classes.

You built a model using Lasso regularization but some important features were wrongly set to zero. How would you modify your approach to keep these features?

  • Combine with ElasticNet
  • Decrease L1 penalty
  • Increase L1 penalty
  • Switch to Ridge
Combining with ElasticNet allows for balancing between L1 and L2 penalties, thus avoiding complete elimination of important features by the L1 penalty.

If multicollinearity is a concern, ________ regularization can provide a solution by shrinking the coefficients.

  • ElasticNet
  • Lasso
  • Ridge
  • nan
Ridge regularization provides a solution to multicollinearity by shrinking the coefficients through the L2 penalty, which helps to stabilize the estimates.

What are some common methods to detect multicollinearity in a dataset?

  • Adding more data
  • Feature scaling
  • Regularization techniques
  • VIF, Correlation Matrix
Common methods to detect multicollinearity include calculating the Variance Inflation Factor (VIF) and examining the correlation matrix among variables.

Name a popular algorithm used in classification problems.

  • Clustering
  • Decision Trees
  • Linear Regression
  • Principal Component Analysis
Decision Trees are a popular algorithm used in classification problems. They work by recursively partitioning the data into subsets based on feature values, leading to a decision on the class label.