Unit 2: Data analysis - Practice Quiz

BTY587 — Data Analysis And Simulations 60 Questions
0 Correct 0 Wrong 60 Left
0/60

1 What does descriptive statistics primarily aim to do?

statistical analysis Easy
A. Prove that a theory is universally true
B. Generate random numbers for simulations
C. Encrypt data for secure storage
D. Summarize and describe the main features of a dataset

2 Which measure of central tendency is most affected by extreme outliers?

statistical analysis Easy
A. Mode
B. Range
C. Mean
D. Median

3 The standard deviation is a measure of:

statistical analysis Easy
A. The total sum of all data points
B. The middle value of the data
C. The most frequent value in the data
D. The spread or dispersion of data around the mean

4 In hypothesis testing, the null hypothesis () typically represents:

hypothesis testing Easy
A. A guaranteed positive result
B. No effect or no difference
C. The largest possible effect size
D. The alternative claim being tested

5 A Type I error occurs when we:

hypothesis testing Easy
A. Reject a true null hypothesis
B. Reject a false null hypothesis
C. Fail to collect enough data
D. Accept a true null hypothesis

6 The significance level of a test is usually denoted by which symbol?

hypothesis testing Easy
A.
B.
C.
D.

7 A p-value represents the probability of:

significance of p-value Easy
A. The null hypothesis being true
B. The alternative hypothesis being true
C. Observing results as extreme as the data, assuming is true
D. Making a correct decision in the test

8 If a p-value is less than the significance level , we typically:

significance of p-value Easy
A. Change the alternative hypothesis
B. Accept the null hypothesis
C. Reject the null hypothesis
D. Increase the sample size

9 A smaller p-value indicates:

significance of p-value Easy
A. A larger sample was required
B. Stronger evidence for the null hypothesis
C. A guaranteed real effect exists
D. Stronger evidence against the null hypothesis

10 The chi-square test is most commonly used to analyze:

chi-square Easy
A. Correlation between two continuous variables
B. Continuous numeric averages
C. Categorical (frequency) data
D. Time-series trends

11 A chi-square test of independence checks whether:

chi-square Easy
A. Two means are equal
B. A sample follows a normal distribution
C. Three or more group variances differ
D. Two categorical variables are associated

12 In the chi-square statistic , what does represent?

chi-square Easy
A. The observed frequency
B. The overall mean
C. The number of outcomes
D. The expected frequency

13 A t-test is generally used to compare:

t-test Easy
A. The means of one or two groups
B. The frequencies of categorical data
C. The correlation of two datasets
D. The variances of three or more groups

14 Which type of t-test compares the means of two related measurements from the same subjects?

t-test Easy
A. Independent samples t-test
B. Paired t-test
C. Welch's t-test
D. One-sample t-test

15 The t-test is preferred over the z-test mainly when:

t-test Easy
A. The data is purely categorical
B. More than three groups are compared
C. The population is infinitely large
D. The sample size is small and population variance is unknown

16 ANOVA is primarily used to compare:

ANOVA Easy
A. The median of two groups
B. Means of three or more groups
C. The spread of a single dataset
D. Two categorical variables

17 The test statistic used in ANOVA is the:

ANOVA Easy
A. t-statistic
B. Chi-square statistic
C. F-statistic
D. z-statistic

18 What does the null hypothesis of a one-way ANOVA state?

ANOVA Easy
A. The groups are categorical
B. All group variances are zero
C. At least one mean is different
D. All group means are equal

19 Bayes' theorem is used to update the probability of a hypothesis based on:

bayesian probability Easy
A. The number of groups compared
B. The sample size only
C. New evidence or data
D. The significance level

20 In Bayesian probability, the probability assigned before observing any new data is called the:

bayesian probability Easy
A. Posterior probability
B. Prior probability
C. Marginal probability
D. Likelihood

21 A researcher sets a significance level of and obtains a test statistic that leads to a p-value of . What is the correct decision?

hypothesis testing Medium
A. Reject the null hypothesis
B. Increase the sample size before deciding
C. Fail to reject the null hypothesis
D. Accept the null hypothesis as true

22 Which statement best describes the meaning of a p-value in hypothesis testing?

significance of p-value Medium
A. The probability of observing data as extreme as the sample, assuming the null hypothesis is true
B. The probability that the alternative hypothesis is correct
C. The probability that the null hypothesis is true given the data
D. The probability of making a Type II error in the test

23 You want to compare the mean test scores of two independent groups with small samples and unknown, unequal variances. Which test is most appropriate?

t-test Medium
A. One-sample z-test
B. Welch's two-sample t-test
C. Chi-square test of independence
D. Paired t-test

24 A chi-square goodness-of-fit test compares observed frequencies with expected frequencies . What is the test statistic formula?

chi-square Medium
A.
B.
C.
D.

25 When comparing the means of four different groups simultaneously, why is one-way ANOVA preferred over multiple t-tests?

ANOVA Medium
A. It works only when variances are unequal
B. It controls the inflation of the overall Type I error rate
C. It eliminates the need for a null hypothesis
D. It requires fewer assumptions about the data

26 A disease affects of a population. A test is sensitive and specific. Using Bayes' theorem, which quantity is the posterior probability of disease given a positive test?

bayesian probability Medium
A.
B.
C.
D.

27 A test fails to reject the null hypothesis even though the alternative is actually true. What type of error has occurred?

hypothesis testing Medium
A. Sampling error
B. Type I error
C. Standard error
D. Type II error

28 In a study measuring blood pressure of the same patients before and after treatment, which t-test correctly accounts for the data structure?

t-test Medium
A. Paired t-test
B. Independent two-sample t-test
C. Welch's t-test
D. One-way ANOVA

29 A result yields with . Which interpretation is most accurate?

significance of p-value Medium
A. The effect size is guaranteed to be large
B. The result proves the alternative hypothesis is true
C. There is a chance the alternative is correct
D. The result is statistically significant but the effect may still be small

30 A chi-square test of independence uses a contingency table. How many degrees of freedom does it have?

chi-square Medium
A.
B.
C.
D.

31 In one-way ANOVA, what does a large F-statistic generally indicate?

ANOVA Medium
A. The sample sizes are too small to analyze
B. Between-group variance is large relative to within-group variance
C. All group means are exactly equal
D. Within-group variance is large relative to between-group variance

32 In Bayesian inference, how is the posterior distribution related to the prior and the likelihood?

bayesian probability Medium
A. Posterior Likelihood Prior
B. Posterior Prior Likelihood
C. Posterior Likelihood Prior
D. Posterior Prior Likelihood

33 A dataset is strongly right-skewed with several large outliers. Which measure of central tendency best represents its typical value?

statistical analysis Medium
A. Mean
B. Standard deviation
C. Range
D. Median

34 Increasing the sample size in a hypothesis test, while keeping other factors fixed, generally has what effect on statistical power?

hypothesis testing Medium
A. Power becomes undefined
B. Power decreases
C. Power stays exactly the same
D. Power increases

35 A one-sample t-test gives with degrees of freedom and a two-tailed critical value of at . What is the conclusion?

t-test Medium
A. Reject the null hypothesis
B. The test is inconclusive
C. Switch to a chi-square test
D. Fail to reject the null hypothesis

36 In a chi-square test, expected cell counts are recommended to be at least . What is the main reason for this guideline?

chi-square Medium
A. To guarantee the observed counts are whole numbers
B. To make the degrees of freedom larger
C. To force the null hypothesis to be true
D. To ensure the chi-square approximation to the sampling distribution is reliable

37 After a significant one-way ANOVA result, why are post-hoc tests such as Tukey's HSD performed?

ANOVA Medium
A. To confirm the ANOVA assumptions were met
B. To reduce the number of groups compared
C. To identify which specific group means differ from each other
D. To increase the between-group variance

38 A prior probability of rain is . New forecast evidence has likelihood under rain and under no rain. Which value is the denominator (evidence) in Bayes' theorem for updating?

bayesian probability Medium
A.
B.
C.
D.

39 Two variables have a Pearson correlation of . Which conclusion is justified?

statistical analysis Medium
A. One variable definitely causes the other
B. The variables are unrelated
C. The relationship is strongly negative
D. There is a strong positive linear association between the variables

40 A researcher lowers the significance threshold from to . What is the direct effect on the two error types?

significance of p-value Medium
A. Both error probabilities decrease equally
B. Type I error probability decreases while Type II error probability tends to increase
C. Type I error probability increases while Type II decreases
D. Neither error probability is affected

41 A researcher obtains in a two-sided test with and declares the effect "real and important." Which statement is the most accurate critique?

significance of p-value Hard
A. A means there is a probability the null hypothesis is true
B. A -value just below neither quantifies effect size nor the probability the null is true; significance and importance are distinct
C. The result cannot be significant because is too close to the threshold
D. A -value this small guarantees the finding will replicate

42 In a fixed-sample test, holding effect size and sample size constant, what happens to the Type II error rate as you lower from to ?

hypothesis testing Hard
A. is unaffected by
B. becomes equal to
C. increases (power decreases)
D. decreases (power increases)

43 Two groups have equal sample sizes but markedly unequal variances. Using a Student's pooled-variance -test instead of Welch's -test primarily risks:

t-test Hard
A. Making the test invalid only when exceeds 30
B. Always inflating power regardless of variance ratio
C. Producing an identical result to Welch's test in all cases
D. Distorting the Type I error rate because the pooled variance misestimates the standard error

44 A one-way ANOVA across 4 groups yields a significant . What is the correct next step and why?

ANOVA Hard
A. Conclude all four means are pairwise different
B. Run 6 independent -tests at each with no correction
C. Run post-hoc pairwise comparisons with a correction (e.g., Tukey HSD) because ANOVA only shows that at least one mean differs
D. Conclude only the largest and smallest means differ

45 In a chi-square test of independence on a table, several expected cell counts fall below 5. The most appropriate response is to:

chi-square Hard
A. Increase to compensate for the small counts
B. Report the chi-square statistic anyway with no adjustment
C. Use Fisher's exact test, since the chi-square approximation is unreliable with small expected counts
D. Convert counts to proportions and use a -test on raw percentages

46 A disease affects of a population. A test has sensitivity and specificity. Given a positive result, the posterior probability of disease is approximately:

bayesian probability Hard
A.
B.
C.
D.

47 A dataset is strongly right-skewed with several extreme high values. Which pair of summary statistics best describes its center and spread?

statistical analysis Hard
A. Mode and range
B. Mean and standard deviation
C. Median and interquartile range (IQR)
D. Mean and variance

48 Researchers test 100 independent true null hypotheses each at . What is the expected number of false positives, and what does this illustrate?

hypothesis testing Hard
A. About 5 false positives, illustrating the multiple-comparisons problem
B. Exactly 1 false positive, fixed by
C. About 50 false positives, because compounds multiplicatively
D. Zero false positives, because all nulls are true

49 A paired-samples design (before/after on the same subjects) is incorrectly analyzed with an independent two-sample -test. The typical consequence is:

t-test Hard
A. An identical -statistic to the paired test
B. Automatic Type I error inflation regardless of correlation
C. The test becomes invalid only for negative correlations
D. Loss of power because the analysis ignores within-subject correlation and inflates the standard error

50 In a one-way ANOVA, if the between-group variability is small relative to within-group variability, the -statistic will be:

ANOVA Hard
A. Very large, indicating strong group differences
B. Negative, indicating error in computation
C. Exactly 0 whenever means are unequal
D. Close to 1, suggesting group means do not differ beyond chance

51 Which interpretation of a confidence interval that excludes 0 for a mean difference is correct?

significance of p-value Hard
A. The true difference is definitely nonzero
B. There is a probability the true difference lies in this specific interval
C. The -value must be exactly
D. The corresponding two-sided test rejects at

52 A chi-square goodness-of-fit test compares observed counts across 6 categories to a theoretical distribution with no estimated parameters. The degrees of freedom are:

chi-square Hard
A. 5
B. 6
C. 1
D. 4

53 In Bayesian inference, as the sample size grows very large, the influence of a proper prior on the posterior typically:

bayesian probability Hard
A. Diminishes, so the posterior is dominated by the likelihood
B. Grows, so the prior overwhelms the data
C. Stays constant regardless of data
D. Becomes undefined for continuous parameters

54 A strong positive linear correlation () is found between ice-cream sales and drowning incidents. The best explanation is:

statistical analysis Hard
A. Ice-cream consumption causes drowning
B. Drownings cause increased ice-cream sales
C. A likely confounder (e.g., temperature) drives both; correlation does not imply causation
D. The correlation must be a computational artifact

55 For a one-sample -test with , the degrees of freedom and the reason are:

t-test Hard
A. , because two parameters are estimated
B. , because
C. , because equals the sample size
D. , because one parameter (the sample mean) is estimated from the data

56 A study is 'underpowered.' What is the most direct risk this poses to interpreting a non-significant result?

hypothesis testing Hard
A. The confidence interval will be too narrow
B. The Type I error rate is inflated above
C. The -value is guaranteed to be below
D. A true effect may exist but go undetected (high Type II error)

57 A two-way ANOVA shows a significant interaction between Factor A and Factor B. How should the main effects be interpreted?

ANOVA Hard
A. With caution, because the effect of one factor depends on the level of the other
B. As automatically non-significant
C. As fully independent and directly interpretable
D. As identical to the interaction effect

58 Which practice most directly leads to 'p-hacking' and inflated false-positive rates?

significance of p-value Hard
A. Trying many analyses and reporting only those with
B. Pre-registering a single hypothesis before data collection
C. Reporting effect sizes with confidence intervals
D. Correcting for multiple comparisons

59 For a contingency table with rows and columns, the degrees of freedom for the chi-square test of independence are:

chi-square Hard
A.
B.
C.
D.

60 You flip a coin 10 times and get 8 heads. Using a uniform prior on the head probability , the posterior distribution is:

bayesian probability Hard
A.
B.
C.
D.