Explain the difference between hard and soft classification.
- Hard provides class labels; soft provides class probabilities
- Hard requires more data; soft requires less
- Hard uses algorithms; soft uses manual classification
- No difference
Hard classification provides specific class labels, whereas soft classification provides probabilities for each class, allowing for more nuanced insights into the confidence of a prediction.
Imagine you're working on a binary classification problem, and the model is performing well in terms of accuracy but poorly in terms of recall. What might be the issue and how would you address it?
- Issue with data imbalance; Use resampling techniques
- Issue with precision; Improve accuracy
- Threshold is too high; Lower the threshold
- Threshold is too low; Increase the threshold
The issue might be that the threshold for classification is set too high, causing true positives to be misclassified as false negatives, reducing recall. Lowering the threshold may help in improving recall without sacrificing too much precision.
What is the primary challenge in implementing unsupervised learning as compared to supervised learning?
- Difficulty in validation
- Lack of rewards
- Requires more data
- Uses only labeled data
The primary challenge in unsupervised learning is the difficulty in validation since there are no predefined labels to assess the model's accuracy.
What is the effect of increasing the regularization parameter in Ridge and Lasso regression?
- Decrease in bias and increase in variance
- Increase in bias and decrease in variance
- Increase in both bias and variance
- No change in bias and variance
Increasing the regularization parameter leads to greater regularization strength, resulting in an increase in bias and a decrease in variance, thus constraining the model complexity.
How does dimensionality reduction help in reducing the risk of overfitting?
- All of the above
- By reducing noise
- By removing irrelevant features
- By simplifying the model
Dimensionality reduction helps in reducing the risk of overfitting by removing irrelevant features (reducing complexity), reducing noise (avoiding fitting to noise), and simplifying the model (making it more generalized).
You are dealing with a dataset having many irrelevant features. How would you apply Lasso regression to deal with this scenario?
- By increasing the degree of the polynomial
- By using L1 regularization
- By using L2 regularization
- By using both L1 and L2 regularization
Lasso regression applies L1 regularization, which can shrink the coefficients of irrelevant features to exactly zero. This effectively performs feature selection, removing the irrelevant features from the model and simplifying it.
You have a highly imbalanced dataset with rare positive cases. Which performance metric would be the most informative, and why?
- AUC, as it provides a comprehensive evaluation of the model
- Accuracy, as it gives overall performance
- F1-Score, as it balances Precision and Recall
- Precision, as it focuses on false positives
In a highly imbalanced dataset, F1-Score is often most informative as it balances Precision and Recall. Accuracy might be misleading, and while AUC and Precision are useful, F1-Score provides a better overall sense of how well the model handles both classes.
What does Precision measure in classification problems?
- False Positives / Total predictions
- True Negatives / (True Negatives + False Positives)
- True Positives / (True Positives + False Negatives)
- True Positives / (True Positives + False Positives)
Precision is the ratio of true positive predictions to the sum of true positives and false positives. It focuses on the accuracy of the positive predictions and is particularly important when the cost of false positives is high.
The technique called ___________ can be used for nonlinear dimensionality reduction, providing a way to reduce dimensions while preserving the relationships between instances.
- PCA
- clustering
- normalization
- t-SNE
t-SNE (t-distributed Stochastic Neighbor Embedding) is a technique used for nonlinear dimensionality reduction. It's effective at preserving the relationships between instances in the reduced space, making it suitable for complex datasets where linear methods like PCA might fail.
Suppose you have hierarchical data and need to understand the relationships between different parts. How would you approach clustering in this context?
- Use DBSCAN
- Use Hierarchical Clustering
- Use K-Means
- Use Mean Shift
Hierarchical Clustering is well-suited for understanding relationships within hierarchical data, as it creates a tree-like structure representing data hierarchies.