Unit 5: Role of Statistics in Data Science - Practice Quiz

ECAP792 60 Questions
0 Correct 0 Wrong 60 Left
0/60

1 What is the main purpose of hypothesis testing?

Hypothesis testing Easy
A. To create charts from raw data
B. To organize records in a database
C. To remove all variation from data
D. To evaluate a claim using sample data

2 What does the null hypothesis usually state?

Null hypothesis Easy
A. There is a strong positive effect
B. There is no effect or difference
C. Every variable has equal values
D. Every observation is an outlier

3 What does the alternative hypothesis generally propose?

Alternative hypothesis Easy
A. No relationship can ever exist
B. The population size is unknown
C. The sample has no observations
D. An effect or difference exists

4 A result is commonly considered statistically significant when which condition is met?

Statistical significance Easy
A. The p-value is below the significance level
B. The data contains no missing values
C. The mean equals the median exactly
D. The sample size equals the population size

5 What is a Type 1 error?

Type 1 and type 2 errors Easy
A. Accepting a true null hypothesis
B. Recording an incorrect sample size
C. Rejecting a false null hypothesis
D. Rejecting a true null hypothesis

6 What is a Type 2 error?

Type 1 and type 2 errors Easy
A. Rejecting a false null hypothesis
B. Failing to reject a false null hypothesis
C. Rejecting a true null hypothesis
D. Selecting an unsuitable chart type

7 What does a p-value measure?

p-value Easy
A. Evidence against the null hypothesis
B. The accuracy of every observation
C. The number of variables collected
D. The exact size of the population

8 If and , what is the usual decision?

p-value Easy
A. Increase the p-value to
B. Retain every possible hypothesis
C. Reject the null hypothesis
D. Conclude that no data exists

9 What is ANOVA mainly used to compare?

ANOVA Easy
A. Names of three or more variables
B. Colors of two different charts
C. Sizes of two database tables
D. Means of three or more groups

10 What does ANOVA stand for?

ANOVA Easy
A. Assessment of Variable Accuracy
B. Analysis of Visual Attributes
C. Analysis of Variance
D. Arrangement of Variable Averages

11 What does the null hypothesis in a basic one-way ANOVA state?

ANOVA Easy
A. All observations are identical
B. All group means are equal
C. All variables are categorical
D. All group sizes are equal

12 Which type of data is commonly analyzed using a chi-square test?

Chi-square test Easy
A. Unstructured image data
B. Continuous measurement data
C. Ordered time-series data
D. Categorical frequency data

13 A chi-square test of independence examines whether two categorical variables are what?

Chi-square test Easy
A. Normally distributed together
B. Measured in identical units
C. Equal in numerical value
D. Associated with each other

14 How does statistics support data science?

Data science Easy
A. It helps analyze data and draw conclusions
B. It guarantees every prediction is correct
C. It converts all data into images
D. It replaces the need to collect data

15 What symbol commonly represents the significance level?

Statistical significance Easy
A.
B.
C.
D.

16 Which two statements are directly compared in hypothesis testing?

Hypothesis testing Easy
A. Sample and chart titles
B. Null and alternative hypotheses
C. Mean and median values
D. Rows and column labels

17 Which everyday term best describes a Type 1 error?

Type 1 and type 2 errors Easy
A. True negative
B. False positive
C. True positive
D. False negative

18 Which everyday term best describes a Type 2 error?

Type 1 and type 2 errors Easy
A. False positive
B. True positive
C. True negative
D. False negative

19 Which p-value provides stronger evidence against the null hypothesis?

p-value Easy
A.
B.
C.
D.

20 What does the null hypothesis in a chi-square test of independence usually state?

Chi-square test Easy
A. The numerical variables are normally distributed
B. The categorical variables are independent
C. The numerical variables have equal means
D. The categorical variables have equal names

21 A data scientist measures the prediction latency of the same 30 servers before and after an optimization. Which test is most appropriate for comparing the mean latencies?

Hypothesis testing Medium
A. Independent-samples t-test
B. Chi-square independence test
C. One-way ANOVA
D. Paired-samples t-test

22 An online retailer tests whether a redesigned page changes the conversion rate from its current value of 8%. Which null hypothesis is appropriate?

Null hypothesis Medium
A.
B.
C.
D.

23 A company will adopt a new recommendation algorithm only if it increases mean revenue above the current value of per user. Which alternative hypothesis should be used?

Alternative hypothesis Medium
A.
B.
C.
D.

24 A hypothesis test produces a p-value of with significance level . What is the appropriate decision?

Statistical significance Medium
A. Reject as practically significant
B. Retain as statistically significant
C. Increase before deciding
D. Reject as statistically significant

25 In a fraud-detection test, states that a transaction is legitimate. What is a Type I error?

Type 1 and type 2 errors Medium
A. Approving a legitimate transaction as legitimate
B. Flagging a fraudulent transaction as fraud
C. Approving a fraudulent transaction as legitimate
D. Flagging a legitimate transaction as fraud

26 A medical screening study uses : the patient does not have the condition. Which outcome represents a Type II error?

Type 1 and type 2 errors Medium
A. Flagging a patient without the condition
B. Missing a patient who has the condition
C. Correctly identifying a patient with the condition
D. Correctly clearing a patient without the condition

27 A data scientist wants to reduce the probability of a Type II error while keeping fixed. Which change is generally most effective?

Type 1 and type 2 errors Medium
A. Increase the sample size
B. Decrease the sample size
C. Use fewer observations
D. Increase measurement noise

28 A model comparison produces a p-value of . Which interpretation is correct?

p-value Medium
A. The probability that is true is exactly 0.02
B. The observed effect has a 0.02 probability of being useful
C. Assuming is true, equally or more extreme data have probability 0.02
D. The probability that is false is exactly 0.02

29 A test returns . At which significance levels would the result be statistically significant?

p-value Medium
A. and
B. and
C. and
D. and

30 A data scientist wants to compare mean customer spending across four membership tiers. Which test is the most appropriate initial analysis?

ANOVA Medium
A. Pearson correlation test
B. One-way ANOVA
C. Paired-samples t-test
D. Chi-square goodness-of-fit test

31 Which set of conditions best matches the standard assumptions of a one-way ANOVA?

ANOVA Medium
A. Paired observations, uniform residuals, and unequal medians
B. Independent observations, normal residuals, and similar variances
C. Repeated observations, skewed residuals, and unequal ranges
D. Independent categories, binary residuals, and equal frequencies

32 An ANOVA comparing five model configurations is statistically significant. What should the data scientist do next to identify which means differ?

ANOVA Medium
A. Accept that all five means differ
B. Replace every mean with its median
C. Apply a post-hoc multiple-comparison test
D. Run one test on the combined sample

33 A researcher examines whether preferred device type (mobile, tablet, desktop) is associated with subscription status (free, paid). Which test is appropriate?

Chi-square test Medium
A. Paired-samples t-test
B. One-sample t-test
C. One-way repeated-measures ANOVA
D. Chi-square test of independence

34 In a contingency table, a row total is 120, a column total is 50, and the grand total is 200. What is the expected count for their intersecting cell?

Chi-square test Medium
A.
B.
C.
D.

35 A die is rolled 120 times to determine whether its outcomes follow a uniform distribution. Which statistical test should be used?

Chi-square test Medium
A. Chi-square goodness-of-fit test
B. One-way repeated-measures ANOVA
C. Chi-square test of independence
D. Two-sample t-test

36 What are the degrees of freedom for a chi-square test of independence using a contingency table?

Chi-square test Medium
A.
B.
C.
D.

37 A machine-learning model has 92% validation accuracy, while another has 91.8%. Why should a data scientist perform a statistical comparison?

Data science Medium
A. To guarantee the first model will perform best in production
B. To remove the need for an independent test dataset
C. To prove that validation accuracy measures business value
D. To assess whether the difference may reflect sampling variation

38 A team independently tests 10 model features and wants a family-wise significance level of 0.05 using the Bonferroni correction. What threshold should each test use?

Hypothesis testing Medium
A.
B.
C.
D.

39 With a very large dataset, a new model reduces mean error by 0.001 units and produces . What is the best conclusion?

Statistical significance Medium
A. The reduction proves the model will improve every prediction
B. The reduction is statistically significant but may lack practical importance
C. The reduction is not statistically significant because it is small
D. The reduction is practically important because the p-value is small

40 In a one-way ANOVA, what does a large -statistic generally indicate?

ANOVA Medium
A. All group means and variances are necessarily equal
B. Between-group variation is large relative to within-group variation
C. The response and group variables are both categorical
D. Within-group variation is large relative to between-group variation

41 A researcher observes a symmetric test statistic and chooses the one-sided alternative only after seeing whether is positive or negative. They then reject whenever the selected one-sided p-value is below . Under , what is the actual Type I error rate?

Hypothesis testing Hard
A. Approximately
B. Approximately
C. Approximately
D. Approximately

42 For a composite null hypothesis , which p-value construction most directly guarantees that a level- test controls Type I error for every ?

Null hypothesis Hard
A. for unrestricted
B.
C.
D. for null

43 To establish that a mean difference is practically equivalent to zero within a margin , which hypotheses are appropriate for the two one-sided tests procedure?

Alternative hypothesis Hard
A. versus
B. versus
C. versus
D. versus

44 Twenty independent parameters each have a pointwise confidence interval. If all model assumptions hold, what is the probability that all 20 intervals simultaneously contain their parameters?

Statistical significance Hard
A.
B.
C.
D.

45 For a fixed nonzero effect size and correctly specified model, sample size is increased while the test threshold is recalibrated to keep . What generally happens?

Type 1 and type 2 errors Hard
A. Type I error stays fixed and Type II error increases
B. Type I error stays fixed and Type II error decreases
C. Type I error increases and Type II error decreases
D. Type I error decreases and Type II error stays fixed

46 A data scientist performs 20 independent hypothesis tests, each at significance level , and every null hypothesis is true. What is the probability of at least one Type I error?

Type 1 and type 2 errors Hard
A.
B.
C.
D.

47 A valid test produces . Which interpretation is correct?

p-value Hard
A. The probability that is true equals
B. Under , results at least as extreme have probability
C. If repeated, the probability of replication equals
D. The probability that the observed result arose by chance equals

48 Under , a one-sided test rejects for unusually large . If is observed, what is the exact p-value?

p-value Hard
A.
B.
C.
D.

49 In a balanced ANOVA, treatment B exceeds treatment A by units at the low level of a second factor, but treatment B is units below treatment A at the high level. Which conclusion is most defensible?

ANOVA Hard
A. The interaction must vanish because effects have equal magnitude
B. The treatment must have a negative main effect
C. The treatment must have a positive main effect
D. The treatment main effect can vanish despite an interaction

50 Three groups each contain four observations. Their sample means are , , and , while the pooled within-group sum of squares is . What is the one-way ANOVA statistic ?

ANOVA Hard
A.
B.
C.
D.

51 A study records five repeated measurements from each participant but analyzes all measurements using ordinary one-way ANOVA as if they came from different people. With positive within-participant correlation, what is the main inferential risk?

ANOVA Hard
A. Residual variance is necessarily estimated without bias
B. Standard errors are overestimated and Type I error is inflated
C. Group means are necessarily biased toward zero
D. Standard errors are underestimated and Type I error is inflated

52 A one-way comparison has strongly unequal group variances and unequal sample sizes, with the largest variance occurring in the smallest group. Which omnibus procedure is generally most appropriate?

ANOVA Hard
A. Repeated-measures ANOVA using participant-level blocks
B. Welch's ANOVA using variance-adjusted group weights
C. Two-way ANOVA using variance as another factor
D. Classical ANOVA using the pooled within-group variance

53 A chi-square goodness-of-fit test uses six categories. Two distribution parameters are estimated from the same observations before expected counts are calculated. What are the test's degrees of freedom?

Chi-square test Hard
A.
B.
C.
D.

54 A table has observed counts . What is Pearson's chi-square statistic without continuity correction?

Chi-square test Hard
A.
B.
C.
D.

55 Binary outcomes are measured on the same individuals before and after an intervention. Which test appropriately assesses marginal change while accounting for pairing?

Chi-square test Hard
A. One-way ANOVA applied to the four cell frequencies
B. McNemar's chi-square test using discordant pairs
C. Chi-square goodness-of-fit test using diagonal pairs
D. Pearson's chi-square independence test on all four cells

56 A feature-selection algorithm is run once on the complete labeled dataset, after which cross-validation evaluates a model using only the selected features. Why is the estimated performance generally optimistic?

Data science Hard
A. Cross-validation requires every available feature in every model
B. Feature selection always increases the model's irreducible error
C. Validation labels influenced selection before the folds were evaluated
D. The folds become unequal because selected features change sample counts

57 Suppose outcomes are missing at random conditional on observed covariates , and every observation has a positive response probability. Which method can consistently estimate the population outcome mean if its required model is correctly specified?

Data science Hard
A. Inverse-probability weighting using modeled response probabilities
B. Complete-case averaging without using
C. Deleting covariates associated with the missingness indicator
D. Mean imputation using only observed outcomes

58 Five ordered p-values are . Using the Benjamini–Hochberg procedure with false discovery rate , how many hypotheses are rejected?

Statistical significance Hard
A. Three hypotheses
B. Five hypotheses
C. Two hypotheses
D. Four hypotheses

59 For four equally sized groups, which pair of contrasts is orthogonal?

ANOVA Hard
A. and
B. and
C. and
D. and

60 An aggregated chi-square test shows a strong association between treatment and recovery, but the association largely disappears within every hospital. Hospital is related to both treatment assignment and recovery. What is the best next analysis?

Chi-square test Hard
A. Use a stratified method such as the Cochran–Mantel–Haenszel test
B. Apply a goodness-of-fit test only to the treatment totals
C. Combine sparse hospitals until the aggregate association becomes stronger
D. Repeat the aggregated test with a smaller significance level