Unit 7: Data Collection Methods

DEMGN832 — Research Methodology 11 min read

I. Foundations of Data Collection

Data collection is the systematic process of obtaining observations or measurements relevant to a research problem. It follows the formulation of objectives, research questions, and hypotheses, and it supplies the evidence used for analysis and interpretation.

Defining characteristics:

  • Purposefulness: Every item collected should relate to a research objective, variable, hypothesis, or analytical requirement.
  • Primary data: Information collected directly for the current study, such as interview responses, experimental measurements, or observed behaviour.
  • Secondary data: Existing information originally collected for another purpose, such as census tables, administrative records, journals, or company reports.
  • Quantitative data: Numerical measurements or coded categories suitable for statistical analysis, such as age, income, test scores, or purchase frequency.
  • Qualitative data: Non-numerical material used to understand meanings, experiences, and processes, such as field notes or open-ended responses.
  • Measurement quality:
    • Reliability means that a procedure produces consistent results under comparable conditions.
    • Validity means that the procedure measures what it is intended to measure.
  • Standardisation: Instructions, instruments, timing, and recording rules should be sufficiently consistent to make cases comparable.
  • Ethical control: Researchers must obtain appropriate consent, protect confidentiality, minimise harm, and collect only data necessary for the stated purpose.
  • Method selection: The choice among observation, experimentation, and surveys depends on the research question, required control, population, resources, and acceptable degree of intrusion.

II. Observation — Recording Behaviour and Events

A. Observation methods

Observation methods collect data by systematically watching, listening to, and recording behaviours, events, objects, or conditions as they occur.

  • Direct observation: The researcher records the behaviour itself, such as counting how many customers enter a store between 10:00 and 11:00.
  • Indirect observation: Evidence or traces of past behaviour are examined, such as product wear, website logs, or archival records.
  • Participant observation: The observer joins the group or setting being studied. For example, a researcher may work alongside employees to understand workplace routines.
  • Non-participant observation: The observer remains separate from the activity and records events without joining them.
  • Structured observation: Predetermined categories and recording rules are used. A classroom study might code behaviour as “asking a question,” “taking notes,” or “off-task.”
  • Unstructured observation: The observer records broad field notes without restricting attention to fixed categories, making it useful in exploratory research.
  • Natural observation: Behaviour is studied in its ordinary setting, such as passengers’ use of ticket machines at a railway station.
  • Contrived observation: The researcher creates or arranges the setting, such as inviting users into a laboratory to test a new interface.
  • Disguised observation: Participants do not know the specific behaviour being studied, reducing reactivity but raising consent and privacy concerns.
  • Undisguised observation: Participants know that observation is taking place, which is ethically clearer but may alter their behaviour.
  • Mechanical observation: Devices such as scanners, cameras, eye trackers, traffic counters, and server logs record events automatically.
  • Observation schedule: A structured form specifies the unit of observation, categories, time interval, and counting rule. Categories should be mutually exclusive and clearly defined.
  • Inter-observer reliability: Agreement between observers can be expressed as:
TEXT
Percentage agreement = Agreements / Total observations × 100

Here, agreements are cases coded identically, and total observations are all cases compared.

B. Applications and limitations

Observation is strongest when actual behaviour matters more than participants’ descriptions of their behaviour.

  • Applications: It is used in consumer research, ethnography, usability testing, classroom studies, traffic analysis, and studies of workplace practice.
  • Behavioural accuracy: It avoids errors caused by faulty memory or socially desirable answers because the researcher records what occurs.
  • Contextual value: Field observation can show where, when, and under what conditions an action takes place.
  • Limited access: Private actions, attitudes, intentions, and motives cannot usually be observed directly.
  • Observer effect: People may change their behaviour when they know they are being watched; this is often called reactivity.
  • Observer bias: Expectations may influence what is noticed or recorded, so training and explicit coding rules are necessary.
  • Time and ethics: Extended observation can be expensive, while covert recording may violate privacy or informed-consent requirements.

III. Experimentation — Testing Causal Relationships

A. Experimentation methods

Experimentation methods examine causality by manipulating an independent variable and observing its effect on a dependent variable while controlling alternative explanations.

  • Independent variable: The factor deliberately changed by the researcher, such as an advertisement’s price message.
  • Dependent variable: The measured outcome, such as purchase intention or sales volume.
  • Experimental group: Participants receive the treatment or intervention being tested.
  • Control group: Participants do not receive the treatment, or receive a standard condition, providing a basis for comparison.
  • Random assignment: Participants are allocated to conditions by chance, helping distribute pre-existing differences across groups.
  • Extraneous variable: Any outside factor that could affect the dependent variable, such as prior brand familiarity.
  • Confounding variable: An uncontrolled factor that changes systematically with the independent variable and creates a rival explanation.
  • Laboratory experiment: Conducted in a controlled environment, offering strong control but sometimes producing artificial behaviour.
  • Field experiment: Conducted in a natural setting, such as testing two shelf displays in actual stores, improving realism but reducing control.
  • Between-subjects design: Different participants experience different experimental conditions.
  • Within-subjects design: The same participants experience multiple conditions; order effects must then be controlled through randomisation or counterbalancing.
  • Pre-test/post-test design: The dependent variable is measured before and after treatment. A basic treatment effect may be represented as:
TEXT
Treatment effect = (E₂ − E₁) − (C₂ − C₁)

E₁ and E₂ are experimental-group pre-test and post-test scores; C₁ and C₂ are corresponding control-group scores.

B. Applications and limitations

Experiments provide the strongest evidence of cause and effect when manipulation, temporal order, and control are successfully established.

  • Causal inference: A credible experiment shows that the treatment preceded the outcome and that plausible rival causes were controlled.
  • Internal validity: This is the degree to which observed changes can be attributed to the treatment rather than history, maturation, testing, or selection effects.
  • External validity: This concerns whether results generalise to other people, places, and times.
  • Manipulation check: A separate measure confirms whether participants noticed or experienced the intended treatment difference.
  • Practical limits: Some variables, such as age or social class, cannot be manipulated, and large field experiments may be costly.
  • Ethical limits: Researchers must not expose participants to unjustified harm, coercion, or deceptive procedures without adequate safeguards and debriefing.

IV. Surveys — Standardised Data from Respondents

A. Survey methods

Survey methods obtain self-reported information from a sample through standardised questions administered in written, oral, electronic, or mixed forms.

  • Cross-sectional survey: Data are collected once from a sample, providing a snapshot of characteristics at a particular time.
  • Longitudinal survey: Data are collected repeatedly to study change.
    • Panel studies revisit the same respondents.
    • Trend studies draw different samples from the same population over time.
  • Face-to-face survey: An interviewer administers questions personally, allowing clarification and complex routing but increasing cost and interviewer influence.
  • Telephone survey: Responses can be obtained quickly across wide areas, although call screening and limited interview length may reduce participation.
  • Mail survey: Respondents complete and return a printed instrument; coverage can be broad, but response rates and turnaround times may be poor.
  • Online survey: Web forms provide rapid distribution, automatic data capture, skip logic, and low marginal cost, but may exclude people with limited digital access.
  • Self-administered survey: Respondents read and answer independently, reducing interviewer bias but requiring clear wording and instructions.
  • Interviewer-administered survey: A trained interviewer asks and records answers, making it suitable for respondents with literacy or accessibility needs.
  • Sampling connection: Survey findings generalise only when the sample adequately represents the target population.
  • Response rate: Participation is commonly calculated as:
TEXT
Response rate = Completed eligible surveys / Eligible sample contacted × 100
  • Survey error: Total error may arise from sampling variation, undercoverage, non-response, inaccurate answers, interviewer effects, and data-processing mistakes.

B. Applications and limitations

Surveys are appropriate when comparable information is required from many people about their characteristics, opinions, knowledge, or reported behaviour.

  • Breadth: A single instrument can measure multiple variables across a geographically dispersed population.
  • Standardisation: Identical wording and response categories support statistical comparison between respondents.
  • Non-response bias: Results become biased when respondents and non-respondents differ on variables central to the study.
  • Recall error: Participants may inaccurately report distant, routine, or sensitive events.
  • Social-desirability bias: Respondents may report behaviour considered acceptable rather than their actual behaviour.
  • Causal limitation: A cross-sectional association between two survey variables does not by itself establish causation or temporal order.

V. Questionnaires — Instrument Development and Data Preparation

A. Introduction to questionnaires

A questionnaire is a formally arranged set of questions used to obtain and record information from respondents in a consistent manner.

  • Research role: It translates abstract objectives and constructs into questions and measurable responses.
  • Closed-ended questions: Respondents select from predefined alternatives, such as “Yes/No” or a five-point satisfaction scale.
  • Open-ended questions: Respondents answer in their own words, producing richer detail but requiring interpretation and coding.
  • Dichotomous questions: Exactly two alternatives are supplied, such as “Employed: Yes or No.”
  • Multiple-choice questions: Several categories are offered; instructions must state whether one or multiple answers are permitted.
  • Rating scales: Respondents indicate degree or intensity. A five-point item may run from 1 = Strongly disagree to 5 = Strongly agree.
  • Filter questions: These determine whether later questions apply. A non-user of a service may be directed past questions about service satisfaction.
  • Measurement levels: Responses may be nominal categories, ordinal ranks, interval-like scale scores, or ratio measurements such as age in years.

B. Questionnaire design process

The questionnaire design process converts research information needs into a clear, valid, ethical, and administrable measurement instrument.

  • Define objectives: Specify the information required and connect each proposed item to a variable or research question.
  • Identify respondents: Consider their language, literacy, knowledge, accessibility needs, and willingness to disclose information.
  • Choose administration mode: Online, mail, telephone, and face-to-face modes impose different limits on question length, visual aids, and routing.
  • Select question content: Include only necessary items and avoid requesting information respondents cannot reasonably know or remember.
  • Write clear wording: Use familiar, specific, neutral language and define reference periods, such as “during the past seven days.”
  • Avoid defective questions:
    • Leading questions suggest a preferred response.
    • Double-barrelled questions ask about two matters at once, such as satisfaction with “price and quality.”
    • Loaded questions contain assumptions or emotionally charged terms.
    • Double negatives make the intended meaning difficult to determine.
  • Construct response options: Categories should be exhaustive, mutually exclusive, balanced, and appropriate to the measurement level.
  • Arrange sequence: Begin with relevant, easy questions; group related items; place sensitive and classification questions later where practical.
  • Design layout: Use consistent numbering, readable spacing, visible instructions, and unambiguous skip directions.
  • Pre-test and revise: Pilot the instrument with people resembling the target population, checking interpretation, timing, routing, missing options, and technical operation.

C. Coding the questionnaire

Coding the questionnaire assigns symbols, usually numbers, to responses so that they can be entered, checked, and analysed systematically.

  • Codebook: A codebook records each variable’s name, description, permissible values, measurement level, and missing-value rules.
  • Closed-response coding: Categories receive unique values, for example 1 = Yes, 2 = No, and 9 = No response.
  • Scale coding: Ordered responses should preserve direction, such as 1 = Very dissatisfied through 5 = Very satisfied.
  • Multiple-response coding: Each option is normally represented by a separate binary variable, where 1 = selected and 0 = not selected.
  • Open-response coding: Researchers review answers, develop categories, define inclusion rules, and test agreement between coders before full coding.
  • Reverse coding: Negatively worded items are recoded before combining scale scores. For a five-point scale:
TEXT
Reversed score = 6 − Original score

Thus, an original score of 2 becomes 4.

  • Missing data: Distinct codes may separate “not applicable,” “refused,” and “not answered,” but these codes must not be treated as substantive numerical values.
  • Data checks: Range checks, skip-pattern checks, and consistency checks identify impossible values, such as an inapplicable follow-up answer.
  • Documentation: Coding decisions must remain stable and traceable so that data cleaning, analysis, and replication use the same definitions.