Unit 3: SUPERVISED LEARNING: CLASSIFICATION - Practice Quiz

INT234 — Predictive Analytics 50 Questions
0 Correct 0 Wrong 50 Left
0/50

1 Which of the following algorithms is categorized as a 'Lazy Learner'?

A. K-Nearest Neighbors
B. Decision Trees
C. Naïve Bayes
D. Support Vector Machines

2 In K-Nearest Neighbors, what is the likely effect of choosing a very small value for 'k' (e.g., k=1)?

A. The model becomes too simple
B. High variance and overfitting
C. High bias and low variance
D. The model ignores local patterns

3 Which distance metric is most commonly used in k-NN for continuous numerical variables?

A. Hamming distance
B. Jaccard similarity
C. Cosine similarity
D. Euclidean distance

4 Why is feature scaling (normalization/standardization) important in k-NN?

A. To increase the value of k
B. To reduce the size of the dataset
C. To convert categorical data to numerical
D. To prevent features with larger scales from dominating the distance calculation

5 The 'Naïve' in Naïve Bayes refers to which fundamental assumption?

A. The prior probabilities are equal
B. All features are dependent on each other
C. The algorithm is simple to implement
D. All features are conditionally independent given the class

6 Naïve Bayes is based on which mathematical theorem?

A. Bayes' Theorem
B. Taylor's Theorem
C. Central Limit Theorem
D. Pythagorean Theorem

7 What is the purpose of Laplace Smoothing in Naïve Bayes?

A. To handle missing values
B. To normalize the data
C. To reduce the dimensionality
D. To handle the problem of zero probability for unseen features

8 Which strategy is primarily used to build Decision Trees?

A. Lazy learning
B. Divide and Conquer
C. Gradient Descent
D. Backpropagation

9 In a Decision Tree, what does a leaf node represent?

A. The root of the tree
B. The entropy value
C. A feature to split on
D. A class label or decision

10 Which metric is commonly used to measure impurity in Decision Trees?

A. Correlation Coefficient
B. Euclidean distance
C. Gini Index
D. R-squared

11 The process of removing branches from a decision tree to prevent overfitting is called:

A. Pruning
B. Regularization
C. Boosting
D. Scaling

12 Which concept represents the expected reduction in entropy caused by partitioning the examples according to an attribute?

A. Maximum Margin
B. Log Loss
C. Gini Impurity
D. Information Gain

13 In rule-based classification, what does 'coverage' refer to?

A. The number of features used
B. The complexity of the rule
C. The accuracy of the rule
D. The number of instances that satisfy the rule's condition

14 The OneR (One Rule) algorithm generates rules based on:

A. All attributes simultaneously
B. A random attribute
C. The single most informative attribute
D. The nearest neighbors

15 What is the primary objective of a Support Vector Machine (SVM)?

A. Minimize the number of features
B. Maximize the posterior probability
C. Find a hyperplane that maximizes the margin between classes
D. Create the deepest possible decision tree

16 The data points that lie closest to the decision boundary in an SVM are known as:

A. Support Vectors
B. Centroids
C. Noise
D. Outliers

17 What technique does SVM use to handle non-linearly separable data?

A. Smoothing
B. Kernel Trick
C. Pruning
D. Bagging

18 In SVM, what is the role of the 'C' hyperparameter?

A. It calculates the Euclidean distance
B. It controls the trade-off between maximizing the margin and minimizing classification errors
C. It sets the depth of the tree
D. It determines the number of kernels

19 In a Confusion Matrix, what does 'False Positive' (Type I Error) represent?

A. Correctly predicting the positive class
B. Incorrectly predicting the negative class when it is actually positive
C. Incorrectly predicting the positive class when it is actually negative
D. Correctly predicting the negative class

20 Which formula correctly calculates Accuracy?

A. (TP + TN) / (TP + TN + FP + FN)
B. TP / (TP + FN)
C. TP / (TP + FP)
D. 2 (Precision Recall) / (Precision + Recall)

21 Accuracy is often a misleading metric when:

A. The dataset is imbalanced
B. The model is a decision tree
C. The dataset is small
D. The dataset is perfectly balanced

22 Which metric represents the ratio of correctly predicted positive observations to the total predicted positives?

A. Recall
B. Precision
C. Specificity
D. Accuracy

23 Recall is also known as:

A. Specificity
B. Precision
C. F1 Score
D. Sensitivity

24 The F1 Score is the harmonic mean of which two metrics?

A. Accuracy and Error Rate
B. True Positive Rate and False Positive Rate
C. Sensitivity and Specificity
D. Precision and Recall

25 Which metric would be most important for a spam detection system where it is acceptable to miss some spam, but critical not to delete legitimate emails (high cost of False Positive)?

A. Sensitivity
B. Log Loss
C. Recall
D. Precision

26 Which metric would be most important for cancer detection where missing a positive case is dangerous (high cost of False Negative)?

A. Specificity
B. Recall
C. Accuracy
D. Precision

27 What does AUC stand for in the context of model evaluation?

A. Average Unit Cost
B. Algorithm User Context
C. Area Under the Curve
D. Accuracy Under Classification

28 The ROC curve plots which two metrics against each other?

A. True Positive Rate vs False Positive Rate
B. Accuracy vs Loss
C. Precision vs Recall
D. Sensitivity vs Specificity

29 An AUC score of 0.5 indicates:

A. A model with high precision
B. A perfect model
C. A model that predicts randomly
D. A model with zero error

30 Logarithmic Loss (Log Loss) penalizes a classifier based on:

A. The confidence of the predicted probabilities
B. The number of misclassifications only
C. The depth of the tree
D. The number of support vectors

31 What is the ideal value for Logarithmic Loss?

A. 0.5
B. 0
C. 1
D. 100

32 Which of the following is a disadvantage of Decision Trees?

A. Cannot handle categorical data
B. Difficult to interpret
C. Prone to overfitting if not pruned
D. Requires feature scaling

33 Which algorithm is generally considered a 'Black Box' model due to low interpretability?

A. Rules (RIPPER)
B. Support Vector Machines (with RBF kernel)
C. Decision Trees
D. Linear Regression

34 In Naïve Bayes, what is the 'Posterior Probability'?

A. The probability of the evidence given the class
B. The probability of the evidence regardless of class
C. The initial probability of the class
D. The probability of the class given the evidence

35 Which classification algorithm is parametric?

A. None of the above
B. K-Nearest Neighbors
C. Naïve Bayes
D. Decision Trees

36 What is the 'Hinge Loss' function associated with?

A. K-Means
B. Logistic Regression
C. Decision Trees
D. Support Vector Machines

37 What happens to the computational cost of k-NN during the prediction phase as the dataset size grows?

A. It becomes zero
B. It increases significantly
C. It remains constant
D. It decreases

38 Entropy in Information Theory is a measure of:

A. Disorder or Uncertainty
B. Distance
C. Accuracy
D. Margin width

39 RIPPER (Repeated Incremental Pruning to Produce Error Reduction) is an algorithm used for:

A. Rule Induction
B. Regression
C. Clustering
D. Dimensionality Reduction

40 Which of the following describes a 'False Negative' (Type II Error)?

A. Predicting Positive when actually Negative
B. Predicting Negative when actually Positive
C. Predicting Positive when actually Positive
D. Predicting Negative when actually Negative

41 If Precision = 1.0 and Recall = 1.0, what is the F1 Score?

A. 2.0
B. 1.0
C. 0.5
D. 0

42 Which evaluation metric calculates the proportion of actual negatives that are correctly identified?

A. Sensitivity
B. Recall
C. Specificity
D. Precision

43 In a decision tree, if a node contains only samples from a single class, its entropy is:

A. 0.5
B. 0
C. 1
D. Infinite

44 Which kernel is the default for non-linear SVMs in many libraries?

A. Radial Basis Function (RBF)
B. Polynomial
C. Sigmoid
D. Linear

45 Generative models like Naïve Bayes model:

A. The boundary between classes directly
B. The error gradients
C. The distribution of individual classes (Joint probability)
D. The distance between points

46 Recursive Partitioning is a technique synonymous with:

A. Calculating k-NN distances
B. Optimizing SVM margins
C. Calculating Bayes probabilities
D. Building Decision Trees

47 When interpreting a Confusion Matrix for a multi-class problem (e.g., 3 classes), the matrix dimensions are:

A. 2x2
B. 3x1
C. 1x3
D. 3x3

48 Which algorithm is most sensitive to outliers?

A. Rules
B. Naïve Bayes
C. Decision Trees
D. K-Nearest Neighbors

49 What is the relationship between Error Rate and Accuracy?

A. Error Rate = Accuracy
B. Error Rate = 1 - Accuracy
C. Error Rate = 1 + Accuracy
D. Error Rate = Accuracy / 2

50 The 'Zero Frequency' problem in Naïve Bayes is solved using:

A. Feature Scaling
B. Laplace Smoothing
C. Pruning
D. Kernel Trick