Unit 6: Measurement and Scaling Technique
I. Orientation
Measurement in research is the systematic assignment of numbers or symbols to characteristics of objects, persons, events, or responses according to clearly defined rules. Scaling extends measurement by placing observations on a continuum that represents the intensity, direction, or amount of a characteristic. In social science, concepts such as attitude, satisfaction, motivation, and loyalty are usually abstract, so researchers measure them through observable indicators.
- Objectivity: The same rules should produce comparable results when applied by different researchers.
- Operational definition: An abstract concept must be translated into observable variables; for example, “academic performance” may be represented by examination marks.
- Standardization: Instructions, response categories, scoring procedures, and administration conditions should remain consistent.
- Validity: The instrument must measure the characteristic it is intended to measure.
- Reliability: The instrument should produce stable and consistent results under similar conditions.
- Levels of measurement: Data may be nominal, ordinal, interval, or ratio, and the permissible statistical operations depend on the level.
- Error control: Observed scores may contain systematic and random error.
II. Tools of Sound Measurement
A. Tools of sound measurement
Tools of sound measurement are instruments and procedures that convert research concepts into dependable observations and numerical data.
- Questionnaire: A structured set of written questions used to collect standardized responses from many respondents. Questions may be open-ended, closed-ended, dichotomous, or multiple-choice.
- Interview schedule: A list of questions administered by an interviewer. It permits clarification of questions but may introduce interviewer bias.
- Observation checklist: A prepared list used to record whether specified behaviours or events occur. For example, a classroom checklist may record “teacher asks questions” or “students participate.”
- Rating scale: A device that records the degree or intensity of a response, such as satisfaction from 1 = very dissatisfied to 5 = very satisfied.
- Attitude scale: A collection of statements designed to measure favourable or unfavourable feelings toward an object, issue, or group.
- Achievement or ability test: A standardized instrument with predetermined scoring rules, often used to assess knowledge, skills, aptitude, or performance.
- Physiological and mechanical devices: Instruments such as weighing scales, stopwatches, blood-pressure monitors, or eye-tracking devices provide physical measurements with defined units.
- Interview and observation protocols: Written instructions improve consistency by specifying the order of questions, observation period, recording method, and treatment of unusual responses.
- Instrument quality: A sound tool should be valid, reliable, sensitive to relevant differences, economical, easy to administer, and acceptable to respondents.
- Sources of measurement error: Ambiguous wording, respondent fatigue, social desirability, interviewer influence, faulty equipment, and inconsistent coding can distort results.
B. Applications and limitations
The usefulness of a measurement tool depends on its fit with the research problem, population, and data-collection setting.
- Application: Questionnaires are efficient for large samples, while interviews are useful when respondents need assistance or when detailed explanations are required.
- Limitation: Self-reports may not accurately reflect actual behaviour because respondents may forget, exaggerate, or provide socially acceptable answers.
- Application: Observation records actual conduct, such as waiting time in a service queue.
- Limitation: The presence of an observer may alter behaviour, and interpretation may differ between observers.
- Practical criterion: A tool requiring 45 minutes per respondent may be unsuitable for a study with 2,000 respondents even if it is highly detailed.
- Ethical criterion: Sensitive questions require privacy, informed consent, and secure handling of personal information.
III. Techniques of Developing Measurement Tools
A. Techniques of developing measurement tools
Developing a measurement tool involves translating a conceptual definition into indicators, items, response formats, and a scoring system.
- Define the construct: State precisely what is being measured and what is excluded. “Job satisfaction” may mean the worker’s positive evaluation of pay, supervision, work, and promotion.
- Identify dimensions: Break a broad construct into components. Customer service quality may include reliability, responsiveness, assurance, and empathy.
- Generate indicators: Use theory, previous instruments, expert opinion, interviews, focus groups, and direct observation to identify observable signs of each dimension.
- Write items: Each item should address one idea, use simple language, avoid leading wording, and specify a suitable time frame. “How satisfied were you with the service received today?” is clearer than “Do you generally like our excellent service?”
- Balance item direction: Positively and negatively worded items may reduce agreement bias, but negative wording must remain clear. Reverse-scored items require careful coding.
- Select response format: A yes/no format is suitable for a factual dichotomy, whereas a five-point agreement scale captures graded attitudes.
- Choose the measurement level: Categories without order are nominal; ordered categories are ordinal; equal numerical intervals support interval measurement; a meaningful zero supports ratio measurement.
- Prepare instructions and scoring: Define how missing answers, multiple marks, reverse items, and total scores will be handled.
- Expert review: Specialists examine content coverage, relevance, clarity, and cultural appropriateness before field administration.
- Pretest or pilot study: Administer the draft to a small group resembling the target population. Record completion time, misunderstood items, nonresponse, and response-pattern problems.
- Item analysis: Examine item difficulty for tests, item-total correlations for scales, and response distributions. Items answered identically by nearly everyone may provide little discrimination.
- Revise and document: Remove redundant or weak items, retain sufficient coverage of the construct, and record the final administration and scoring procedures.
B. Applications and limitations
Tool development improves measurement quality, but statistical refinement cannot compensate for a poorly defined construct.
- Content validity: An employee motivation scale covering only salary omits recognition, autonomy, achievement, and advancement.
- Construct validity: A measure should relate to other variables as theory predicts; motivation should generally show a meaningful relationship with effort or persistence.
- Pilot limitation: A pilot group of 20 university students may not reveal language problems that occur among older rural respondents.
- Cultural limitation: An item about individual career ambition may not have the same meaning in collectivist and individualist settings.
- Respondent burden: Excessively long instruments increase fatigue, straight-line responding, and item nonresponse.
- Documentation: A research report should identify the source of items, modifications, response anchors, scoring direction, and evidence of validity and reliability.
IV. Meaning of Scaling
A. Meaning of scaling
Scaling is the process of assigning numbers to objects or responses so that the numbers represent a meaningful position, category, direction, or intensity on a continuum.
- Purpose: Scaling makes an abstract property measurable; an attitude toward online learning can be represented from strongly negative to strongly positive.
- Continuum: A scale arranges possible positions between two conceptual extremes, such as “very dissatisfied” and “very satisfied.”
- Object assignment: Numbers may describe categories, rank cases, represent equal intervals, or express quantities. Their meaning depends on the rules, not merely on the presence of numbers.
- Item and scale distinction: An individual statement is an item; the combined score from several related items forms a scale.
- Unidimensionality: A scale should ideally measure one underlying attribute. Items about pay, workload, and promotion may need separate dimensions if they do not represent one construct.
- Direction: Researchers must specify whether a high score indicates more of the characteristic. If 1 means “strongly agree” and 5 means “strongly disagree,” interpretation differs from the reverse arrangement.
- Composite score: For five items scored from 1 to 5, a respondent’s total may range from 5 to 25, or the mean may range from 1 to 5.
- Scaling error: Arbitrary or unclear intervals can create false precision; a score of 80 on an attitude scale is not automatically twice as positive as a score of 40.
B. Applications and limitations
Scaling is especially useful for attitudes, preferences, perceptions, and behavioural intentions, but its interpretation must match the scale’s properties.
- Application: A hospital may scale patient satisfaction to compare wards or identify service weaknesses.
- Application: Ranking ten brands reveals preference order but not the distance between first and second choices.
- Limitation: Respondents may avoid extreme categories, agree with statements habitually, or select the middle option without careful consideration.
- Interpretive rule: A numerical code is not automatically a quantity; codes 1, 2, and 3 for departments are nominal labels, not levels of department membership.
V. Important Scaling Techniques
A. Important scaling techniques
Important scaling techniques differ in how they obtain categories, ranks, distances, or agreement levels from respondents.
- Nominal scale: Classifies observations into mutually exclusive categories, such as gender category, blood group, or type of institution. The numbers are labels only.
- Ordinal scale: Arranges observations by rank, such as first, second, and third preference. It shows order but not equal differences.
- Likert summated rating scale: Presents statements with ordered agreement categories, commonly 1 = strongly disagree through 5 = strongly agree. Scores across related items are summed or averaged.
- Semantic differential scale: Uses bipolar adjectives, such as “modern–traditional,” “reliable–unreliable,” or “pleasant–unpleasant,” with several positions between them.
- Stapel scale: Uses one adjective and a numerical range, commonly from +5 to -5; respondents rate an object’s association with the adjective.
- Thurstone equal-appearing interval scale: Experts place attitude statements at presumed scale positions. Respondents indicate agreement, and the researcher uses the assigned median positions to estimate attitude.
- Guttman cumulative scale: Items are arranged from weak to strong expressions of a trait. Agreement with a stronger item should imply agreement with weaker items; the pattern is evaluated for reproducibility.
- Paired comparison scale: Respondents compare two objects at a time, such as Brand A versus Brand B. It is useful for preference ordering but becomes burdensome as the number of objects increases.
- Rank-order scale: Respondents arrange alternatives from most to least preferred. It provides relative priority but does not show how much one alternative exceeds another.
- Constant-sum scale: Respondents distribute a fixed total, often 100 points, among attributes. Giving 40 points to price and 60 to quality expresses relative importance.
B. Applications and limitations
The choice of technique should reflect the research objective, respondent capability, number of alternatives, and required statistical analysis.
- Agreement scales: Likert items efficiently measure attitudes, but acquiescence and central-tendency bias may affect responses.
- Bipolar scales: Semantic differential scales capture the image of a brand, although suitable opposite adjectives can be difficult to identify.
- Cumulative scales: Guttman scaling represents a strong ordering assumption; real responses may violate the expected cumulative pattern.
- Comparative scales: Paired comparisons force a choice between alternatives but require many comparisons. For (n) objects, the number of pairs is:
n(n - 1) / 2Here, (n) is the number of objects. Ten objects therefore require (10(9)/2 = 45) comparisons.
- Point-allocation scales: Constant-sum scoring indicates relative importance, but respondents may find it difficult to allocate points consistently.
- Noncomparative scales: Rating each object independently permits broader comparisons, but different respondents may use the response categories differently.
VI. Statistical Properties of Different Scales
A. Statistical properties of different scales
The statistical properties of a scale determine which summaries and tests are logically defensible.
- Nominal scale: The valid operations are counting frequencies, calculating percentages, identifying the mode, and testing association. The mean of department codes has no substantive meaning.
- Ordinal scale: Median, percentiles, rank-order correlation, and nonparametric tests are appropriate because order is known but intervals are not guaranteed equal.
- Interval scale: Mean, standard deviation, Pearson correlation, regression, and many parametric tests are appropriate when equal intervals are defensible. Celsius temperature illustrates equal intervals without a true zero.
- Ratio scale: All arithmetic operations are meaningful because the scale has equal units and an absolute zero. Weight, income, distance, and response time permit statements such as 80 kg being twice 40 kg.
- Identity: Nominal measurement supports determining whether observations are the same or different.
- Order: Ordinal measurement adds greater-than and less-than relationships.
- Equal intervals: Interval measurement makes differences comparable; the difference between 10 and 20 equals the difference between 40 and 50.
- Absolute zero: Ratio measurement permits meaningful ratios because zero represents absence of the measured quantity.
- Reliability statistics: Internal consistency may be assessed using Cronbach’s alpha; test-retest reliability examines stability over time, while inter-rater reliability examines agreement between observers.
- Validity evidence: Face and content validity concern apparent and substantive coverage; criterion-related validity compares scores with an external criterion; construct validity examines convergent and discriminant relationships.
- Distributional consideration: A five-category ordinal response may be summarized with percentages and medians. Treating a multi-item Likert composite as approximately interval is common when items are coherent and the composite has several categories, but the assumption should be justified.
- Transformation rule: Valid statistical transformations depend on scale type. Nominal labels may be replaced by any other labels, while ratio values may be multiplied by a constant without changing their essential meaning.
- Scale selection principle: The highest defensible level of measurement should be retained because reducing ratio data to categories discards information, whereas claiming interval or ratio properties without justification produces misleading analysis.
Did this save you a night before the exam?
LPU Notes is free, and it stays free. Ads cover part of the server bill. The rest comes out of a student's own pocket: the domain, the storage, and keeping the site up through the weeks everyone needs it at once.
The payment button didn't load. An ad blocker or a filtered network is the usual reason. to try again.
Nothing here is ever locked, and nothing unlocks. Chip in only if it was worth it. What it pays for →