Start by writing what you intend the questionnaire scores to mean and what decision they will support. A survey used to improve a workshop needs different evidence from a scale used to compare organizations or make decisions about individual people.
Validation is an investigation of those interpretations and uses. A published instrument is a useful starting point, but its previous evidence may concern a different population, language, setting, or decision. A high reliability coefficient does not resolve that gap.
Free resource: Open the measurement-review checklist to organize your proposed use, instrument version, available evidence, and next steps. It includes a fictional example and a printable worksheet.
Try completing this sentence: “We will use these scores to describe ___ among ___, so that we can decide ___.” Then identify the consequences of being wrong.
For example, a fictional training organization might want to learn whether participants feel ready to use a skill after a workshop. Self-reported readiness could help identify support needs. It would not, by itself, demonstrate competence or establish that the workshop caused improvement.
Make that distinction explicit in the plan. If the intended claim is about performance, consider whether you also need a task, observation, or other evidence aligned with that claim.
Document the exact version: items, response options, language, recall period, instructions, administration method, and scoring rules. Check what the original studies actually support and how closely their respondents and uses match yours.
Separate a factual questionnaire from a multi-item scale. Questions about attendance, service access, and satisfaction do not automatically form one underlying construct. Combining them into a total score needs a justification; running factor analysis on every survey is not a default requirement.
If you plan to shorten or adapt an instrument, read Can You Change a Validated Questionnaire? before assuming the original evidence will transfer.
Ask whether the items represent the concept you need, leave out an important part, or introduce material unrelated to the intended score. Involve people who understand the subject and people from the respondent population.
A readiness scale that asks only about confidence may miss practical barriers such as access to equipment or supervision. Adding those barriers to the same score is not automatically the answer: they may need separate reporting because they describe a different problem.
Before a large launch, examine whether respondents interpret the questions, recall periods, and response categories as intended. Cognitive interviewing explores the reasoning behind an answer rather than simply asking whether the survey looks clear.
Test the intended format too. A question that works in an interviewer-led session may be misunderstood on a phone screen. Record the problems, revisions, and reasons for keeping or changing an item.
The analyses should follow the proposed score and use. For a multi-item scale, the plan may examine dimensionality, reliability, and relationships with other relevant measures. Exploratory and confirmatory factor analysis answer different questions; repeatedly changing a model to fit one dataset does not provide an independent confirmation.
Choose a sample that supports the planned model and decisions. There is no single sample-size rule that makes every questionnaire valid. Item properties, model complexity, group comparisons, precision, and missing information affect the requirements.
Reliability concerns consistency or precision. It cannot establish what a score means. Do not treat an alpha threshold, one fit index, or a statistically significant correlation as a pass/fail certificate for the entire instrument.
If you want to compare languages, departments, demographic groups, or time points, plan evidence that the scores operate comparably for those comparisons. A difference in average scores may reflect differences in how questions function as well as differences in the concept of interest.
Likewise, evidence suitable for group summaries may not justify individual classifications. A cutoff needs support for its particular decision; it is not established simply by dividing the observed scores into convenient bands.
Ask for a documented instrument version, proposed interpretations, evidence reviewed, analyses conducted, uncertainties, and recommended uses and limits. The result may be a revision plan or a restricted use rather than an unconditional approval.
Bring a blank questionnaire, scoring instructions, population description, and intended decisions to an initial consultation. A data dictionary can describe existing records without sending identifying respondent data through an inquiry form.
DASS can help scope a questionnaire and measurement review, including the evidence needed before you collect another round of responses.
One email whenever we publish something new. No spam, unsubscribe anytime.
{{ subscribeErrorMsg }}