Unit 2: Introduction to Statistics and Data Analysis - Practice Quiz

ECAP790 60 Questions
0 Correct 0 Wrong 60 Left
0/60

1 What is the main purpose of statistical inference?

Statistical Inference Easy
A. To remove all variation from collected data
B. To display every value in a population
C. To arrange observations in alphabetical order
D. To draw conclusions about a population from a sample

2 Which activity is an example of statistical inference?

Statistical Inference Easy
A. Listing the names of all survey participants
B. Sorting survey responses from lowest to highest
C. Estimating voter support using a survey sample
D. Calculating the exact size of a surveyed group

3 In statistics, what is a population?

Samples, Populations and Experimental Design Easy
A. The complete group being studied
B. The average value in a dataset
C. A small part of the collected data
D. A graph of the collected observations

4 Why is random assignment used in an experiment?

Samples, Populations and Experimental Design Easy
A. To create comparable treatment groups
B. To increase the population size
C. To guarantee the expected result
D. To eliminate the response variable

5 What is the sample mean of , , and ?

Measures of Location: The Sample Mean and Median Easy
A.
B.
C.
D.

6 What is the median of the ordered values ?

Measures of Location: The Sample Mean and Median Easy
A.
B.
C.
D.

7 What does the range of a dataset measure?

Measures of Variability Easy
A. The sum of all values divided by their number
B. The value appearing most often in the dataset
C. The middle value after ordering the observations
D. The difference between its largest and smallest values

8 Which measure describes how spread out observations are around their mean?

Measures of Variability Easy
A. Sample mode
B. Standard deviation
C. Sample median
D. Arithmetic mean

9 Which variable is an example of discrete data?

Discrete and Continuous Data Easy
A. Temperature of a classroom
B. Height of a sunflower
C. Time required to finish a race
D. Number of books on a shelf

10 Which variable is an example of continuous data?

Discrete and Continuous Data Easy
A. Count of defective bulbs
B. Weight of a package
C. Number of school buses
D. Number of phone calls

11 What is a statistical model?

Statistical Modeling Easy
A. A mathematical representation of a data-generating process
B. A complete list of observations in alphabetical order
C. A device used to measure physical quantities
D. A rule requiring every observation to be identical

12 In a simple linear model, what does a line commonly describe?

Statistical Modeling Easy
A. The labels assigned to four categories
B. The relationship between two variables
C. The difference between two sample sizes
D. The exact order of all observations

13 What should a researcher do first when scientifically inspecting a dataset?

Scientific Inspection Easy
A. Delete its unusual observations
B. Assume its measurements are correct
C. Replace its missing values with zero
D. Check its values and structure

14 Why should units of measurement be checked during data inspection?

Scientific Inspection Easy
A. To remove the need for data collection
B. To make every observation numerically equal
C. To ensure every variable is categorical
D. To interpret and compare values correctly

15 Which graph is commonly used to inspect residuals from a statistical model?

Graphical Diagnostics Easy
A. Pie chart
B. Frequency table
C. Residual plot
D. Bar chart

16 What does a clear curved pattern in a residual plot often suggest?

Graphical Diagnostics Easy
A. The sample contains only categorical values
B. The fitted relationship may not be linear
C. The fitted model has perfect agreement
D. The response variable has no observations

17 Which graph is most suitable for displaying the distribution of one quantitative variable?

Graphical Methods and Data Description Easy
A. Organization chart
B. Pie chart
C. Flowchart
D. Histogram

18 Which graph is commonly used to compare frequencies across categories?

Graphical Methods and Data Description Easy
A. Residual plot
B. Scatterplot
C. Bar chart
D. Histogram

19 In which type of study does a researcher assign treatments to participants?

General Types of Statistical Studies Easy
A. Census report
B. Case summary
C. Observational study
D. Experimental study

20 What distinguishes an observational study from an experiment?

General Types of Statistical Studies Easy
A. Participants are always selected from one location
B. Every member of the population is measured
C. Treatments are not assigned by the researcher
D. All variables are required to be quantitative

21 A random sample has mean and standard error . Using the approximation estimate standard errors, what is an approximate 95% confidence interval for the population mean?

Statistical Inference Medium
A.
B.
C.
D.

22 A hypothesis test produces a p-value of . Which interpretation is correct?

Statistical Inference Medium
A. Repeating the study will produce the same result with probability .
B. Assuming the null hypothesis is true, results at least this extreme occur with probability .
C. The alternative hypothesis has a probability of of being true.
D. The null hypothesis has a probability of of being true.

23 A farmer wants to compare two fertilizers in a field with a known north-to-south fertility gradient. Which design best controls for this gradient?

Samples, Populations and Experimental Design Medium
A. Allow workers to choose which fertilizer to apply to each plot.
B. Apply both fertilizers only to plots near the center of the field.
C. Apply one fertilizer in the north and the other fertilizer in the south.
D. Divide the field into north-to-south blocks and randomize fertilizers within each block.

24 A company has employees distributed among three departments in proportions , , and . For a proportional stratified sample of 200 employees, how many should be selected from the departments?

Samples, Populations and Experimental Design Medium
A. , , and
B. , , and
C. , , and
D. , , and

25 For the data set , what are the sample mean and median?

Measures of Location: The Sample Mean and Median Medium
A. Mean and median
B. Mean and median
C. Mean and median
D. Mean and median

26 In the data set , the value is replaced by . How do the sample mean and median change?

Measures of Location: The Sample Mean and Median Medium
A. The mean increases by , while the median remains unchanged.
B. The mean remains unchanged, while the median increases by .
C. The mean and median both increase by .
D. The mean increases by , while the median remains unchanged.

27 What is the sample variance of the observations ?

Measures of Variability Medium
A.
B.
C.
D.

28 A variable has a standard deviation of . If , what is the standard deviation of ?

Measures of Variability Medium
A.
B.
C.
D.

29 Which variable is discrete even though it may take many possible values?

Discrete and Continuous Data Medium
A. The number of customers entering a store each day
B. The temperature inside a storage room
C. The time required to serve a customer
D. The weight of a package before shipping

30 A machine's actual fill volume is recorded to the nearest milliliter. How should the underlying fill volume and the recorded value be classified?

Discrete and Continuous Data Medium
A. Both the underlying volume and recorded value are continuous.
B. The underlying volume is discrete, but the recorded value is continuous.
C. The underlying volume is continuous, but the recorded value is discrete.
D. Both the underlying volume and recorded value are discrete.

31 A linear model is fitted to predict fuel use from distance traveled. Its residual plot shows a clear U-shaped pattern. What does this indicate?

Statistical Modeling Medium
A. The linear model misses a nonlinear relationship.
B. The observations are necessarily measured incorrectly.
C. The response has no relationship with distance.
D. The fitted line has exactly the correct form.

32 A regression model was developed using machine speeds from to rpm. It is used to predict output at rpm. What is the main concern?

Statistical Modeling Medium
A. The prediction is an interpolation within the observed range.
B. The response automatically becomes discrete above rpm.
C. The prediction is an extrapolation beyond the observed range.
D. The fitted model cannot calculate values above the sample mean.

33 A laboratory data set contains the measurements , , , and . What is the most appropriate initial action?

Scientific Inspection Medium
A. Keep without investigating because all observations are valid.
B. Check instruments and source records for a measurement or transcription error.
C. Replace with the mean of the remaining measurements.
D. Delete immediately because it differs from the other values.

34 Measurements from a process shift upward immediately after an instrument calibration. Which response best reflects scientific inspection?

Scientific Inspection Medium
A. Assume the shift is random because calibrated instruments cannot create bias.
B. Investigate the calibration and process records before assigning a cause.
C. Conclude that calibration caused the shift because the events were consecutive.
D. Remove all measurements collected after the calibration from analysis.

35 In a residual-versus-fitted plot, residuals are tightly clustered at low fitted values but spread out at high fitted values. What problem is suggested?

Graphical Diagnostics Medium
A. Complete response independence
B. Perfect multicollinearity
C. Constant error variance
D. Nonconstant error variance

36 A normal Q-Q plot follows a straight line in the center but bends away strongly at both ends. What is the most reasonable conclusion?

Graphical Diagnostics Medium
A. The distribution's tails differ from normal tails.
B. The fitted model has a perfect linear relationship.
C. The sample has no variation around its mean.
D. The observations have exactly a normal distribution.

37 A histogram of household incomes has one peak and a long right tail. Which relationship between the mean and median is most likely?

Graphical Methods and Data Description Medium
A. The mean must be twice the median.
B. The mean is equal to the median.
C. The mean is greater than the median.
D. The mean is less than the median.

38 A researcher wants to examine the relationship between hours studied and examination score for 80 students. Which graph is most appropriate?

Graphical Methods and Data Description Medium
A. A scatterplot
B. A frequency polygon
C. A single boxplot
D. A pie chart

39 Researchers compare disease rates among people who voluntarily choose different diets, without assigning the diets. Why is a causal conclusion limited?

General Types of Statistical Studies Medium
A. Causal conclusions require the response variable to be continuous.
B. Voluntary participation guarantees random treatment assignment.
C. Disease rates cannot be measured in observational studies.
D. Diet choice may be associated with confounding variables.

40 Investigators select people with a rare disease and people without it, then examine their past exposure to a chemical. What type of study is this?

General Types of Statistical Studies Medium
A. A randomized controlled experiment
B. A cross-sectional sample survey
C. A retrospective case-control study
D. A prospective cohort study

41 A population is partitioned into strata of known sizes . An independent simple random sample of size is taken from each stratum, producing stratum means . Which estimator is unbiased for the overall population mean when the sampling fractions differ across strata?

Statistical Inference Hard
A.
B.
C.
D.

42 A treatment is randomly assigned to classrooms, but outcomes are recorded for individual pupils. Classrooms have equal sizes, there is no attrition, and outcomes within a classroom are positively correlated. What is the primary consequence of analyzing all pupils as independent observations?

Samples, Populations and Experimental Design Hard
A. The treatment effect is necessarily biased upward, while its standard error remains valid.
B. The effect estimate and standard error remain valid because classroom sizes are equal.
C. The effect estimate can remain unbiased, but its naive standard error is generally too small.
D. The effect estimate becomes inconsistent, but its naive standard error is generally too large.

43 A sample has mean , median , variance , and interquartile range . Every observation is transformed using . Which set of summaries is correct for the transformed sample?

Measures of Location: The Sample Mean and Median Hard
A. Mean , median , variance , and interquartile range
B. Mean , median , variance , and interquartile range
C. Mean , median , variance , and interquartile range
D. Mean , median , variance , and interquartile range

44 Two independent samples have summaries and , where each variance uses denominator . What is the sample variance after combining all observations?

Measures of Variability Hard
A.
B.
C.
D.

45 A variable equals with probability ; otherwise, it follows an exponential distribution with rate . Which description of the distribution of is correct?

Discrete and Continuous Data Hard
A. It is absolutely continuous because its cumulative distribution function is nondecreasing.
B. It is mixed because it has an atom at zero and a continuous component above zero.
C. It is discrete because it assigns positive probability to the value zero.
D. It is continuous because its nonzero values have an exponential density.

46 Suppose , where , , and . Which ordinary least-squares formulation is equivalent to weighted least squares with weights ?

Statistical Modeling Hard
A. Regress on without an intercept and constant error variance.
B. Regress on using an intercept and constant error variance.
C. Regress on without an intercept and constant error variance.
D. Regress on using an intercept and constant error variance.

47 Researchers inspect accumulating results after every 20 observations and stop the first time an ordinary two-sided test gives . Which approach best preserves the intended overall Type I error rate?

Scientific Inspection Hard
A. Use a prespecified group-sequential boundary or an alpha-spending procedure.
B. Report only the final-look p-value because earlier looks were preliminary.
C. Double the final p-value because both directions of the effect were inspected.
D. Retain the cutoff because every individual test has size .

48 A normal Q-Q plot of regression residuals is close to its reference line in the center, falls substantially below the line in the left tail, and rises substantially above it in the right tail. What is the most direct interpretation?

Graphical Diagnostics Hard
A. The response variance decreases steadily as fitted values increase.
B. The residual distribution has lighter tails than the fitted normal distribution.
C. The regression mean function omits a monotone linear component.
D. The residual distribution has heavier tails than the fitted normal distribution.

49 A histogram has bins , , and containing , , and observations, respectively. If the histogram uses density heights so that total area is , which relationship among the bar heights is correct?

Graphical Methods and Data Description Hard
A. The heights have ratio , so the third bar is highest.
B. The heights have ratio , so all three bars have equal area.
C. The heights have ratio , so the third bar is highest.
D. The heights have ratio , so the first bar is highest.

50 Investigators select people with a rare disease and without it, then compare prior exposure histories. With unknown case and control sampling fractions, which effect measure can be estimated directly from this case-control design?

General Types of Statistical Studies Hard
A. The exposure-disease odds ratio based on cases and controls
B. The population disease prevalence among exposed people
C. The population risk difference between exposed and unexposed people
D. The exposed-to-unexposed ratio of disease incidence rates

51 In a paired experiment, each of subjects receives both treatments under randomized order. Which nonparametric bootstrap procedure is appropriate for constructing a confidence interval for the mean within-subject treatment difference?

Statistical Inference Hard
A. Resample the two treatment columns independently and subtract their sample means.
B. Permute treatment labels within every pair and retain the resulting null distribution.
C. Resample the subject-level pairs and recompute the mean paired difference.
D. Pool all measurements, resample them, and divide them into two groups.

52 Units in a finite population of known size are sampled with unequal inclusion probabilities . If is the resulting sample, which estimator is design-unbiased for the population mean?

Samples, Populations and Experimental Design Hard
A.
B.
C.
D.

53 One value is added to the sample . For which value of are the mean and median of the resulting six observations equal?

Measures of Location: The Sample Mean and Median Hard
A.
B.
C.
D.

54 A fixed sample of observations has its largest value replaced by , where tends to infinity and remains the largest value. Which statement about the resulting summaries is correct?

Measures of Variability Hard
A. The mean and median diverge, while the standard deviation remains bounded.
B. The mean and standard deviation diverge, while the median and median absolute deviation remain bounded.
C. The standard deviation and median diverge, while the mean and median absolute deviation remain bounded.
D. The median and median absolute deviation diverge, while the mean remains bounded.

55 A latent measurement is normally distributed, but the instrument reports , the nearest integer to . Ignoring probability-zero boundary ties, what is the correct likelihood contribution from an observed report ?

Discrete and Continuous Data Hard
A. , because integer reports correspond to left-closed unit intervals
B. , because the normal density evaluated at is the probability of observing
C. , because rounding retains only the magnitude of the latent measurement
D. , because all latent values in that interval report

56 A two-component Gaussian mixture is parameterized by . Why can two distinct parameter vectors define exactly the same mixture distribution even with unlimited data?

Statistical Modeling Hard
A. Every Gaussian mixture has an equivalent single-Gaussian representation.
B. Swapping both component labels and replacing by leaves the density unchanged.
C. Mixture weights disappear from the likelihood after observations are marginalized.
D. Gaussian component means cannot be estimated when their variances are unknown.

57 Observed success counts are: mild cases—treatment , control ; severe cases—treatment , control . Which conclusion is supported by inspecting both the aggregate and stratified data?

Scientific Inspection Hard
A. The reported counts are arithmetically incompatible with a reversal after aggregation.
B. Aggregation is harmless because each group contains exactly cases.
C. Treatment has a lower observed success rate within both severity strata.
D. Severity creates a Simpson's-paradox reversal: treatment is higher within each stratum but lower overall.

58 A regression's residual-versus-fitted plot has residuals centered near zero without curvature, but its scale-location plot shows increasing steadily with fitted values. What problem is most strongly indicated?

Graphical Diagnostics Hard
A. The predictors are highly collinear and produce unstable coefficient estimates.
B. The conditional mean function is nonlinear but the variance is constant.
C. The residual distribution is symmetric but has substantially lighter tails.
D. The conditional error variance increases as the fitted response increases.

59 Let be uniformly distributed on and let . Which statement best describes what a scatterplot and the Pearson correlation would show?

Graphical Methods and Data Description Hard
A. The correlation is positive because increases whenever increases.
B. The correlation is undefined because is a deterministic function of .
C. The correlation is negative because negative values of produce positive values of .
D. The correlation is zero, although the scatterplot shows a perfect nonlinear relationship.

60 A difference-in-differences study compares pre-to-post outcome changes in a treated group with those in an untreated group. Which assumption is central for interpreting the difference-in-differences estimate causally?

General Types of Statistical Studies Hard
A. The post-treatment outcomes must have equal variances in the two groups.
B. Before treatment, both groups must have had identical average outcome levels.
C. Treatment assignment must be independent of every measured baseline covariate.
D. Without treatment, the groups would have followed parallel average outcome trends.