Designing a Survey That Produces Valid Data
A survey is only as good as its design. Poorly worded questions, an unrepresentative sample, or unvalidated measures produce data that looks fine and means little—and no amount of sophisticated analysis can rescue it. This guide covers how to design a survey that yields data worth analysing.
Surveys are among the most widely used—and most widely misused—tools in management and social-science research. They look deceptively simple: write some questions, send them out, analyse the answers. But the quality of everything that follows is fixed at the design stage. A survey built on ambiguous questions, a biased sample, or measures that were never validated will produce numbers that can be analysed but cannot be trusted. The most common and costly errors in survey research happen before a single response is collected.
This guide covers the design decisions that determine whether your survey produces valid, reliable data: how to write questions, how to measure constructs, how to sample, and the biases to guard against. It reflects how we approach primary data collection in our Survey & Primary Research practice, and it pairs naturally with our guide to measurement and SEM, since a survey's data usually ends up in exactly those models.
Validity and reliability: the two foundations
Two concepts underpin all survey quality, and they are often confused. Validity is whether your survey actually measures what you intend it to measure. A scale meant to capture “job satisfaction” is valid if it genuinely reflects job satisfaction and not, say, general mood or loyalty to a manager. Reliability is whether it measures consistently—whether the same respondent, in the same state, would give similar answers on repeated administration, and whether items meant to capture the same construct hang together.
The distinction matters because they are independent: a survey can be reliable but not valid (consistently measuring the wrong thing) and, in principle, the reverse. Both are required. A sophisticated structural model built on a measure that is unreliable or invalid is a sophisticated way of being wrong—which is why measurement is assessed before any structural relationship is interpreted. Getting validity and reliability right at the design stage is what makes the eventual analysis meaningful.
Writing questions that work
Most survey damage is done at the level of individual questions. A few disciplines prevent the majority of it.
Ask one thing at a time. Avoid double-barrelled questions—“How satisfied are you with your pay and working conditions?” forces one answer to two different questions. Split them.
Stay neutral. Leading or loaded wording (“How much do you agree that our excellent service…”) pushes respondents toward an answer and biases the data. Word questions so they do not signal a preferred response.
Be clear and concrete. Vague terms, jargon, and ambiguous time frames (“recently,” “often”) mean different things to different people, adding noise. Say what you mean precisely.
Match the response scale to the question, keep scale points balanced (equal positive and negative options), and label them clearly. Decide deliberately whether to offer a neutral midpoint or a “don’t know” option, since each changes how people respond.
Avoid asking what people can’t reliably report. Questions that demand precise recall of distant events, or that ask people to introspect on things they have no access to, produce confident but inaccurate answers.
Pilot the survey before you launch it. A small pilot with people like your target respondents reveals confusing questions, broken logic, and misread scales while you can still fix them. Skipping the pilot is one of the most common—and most avoidable—survey mistakes.
Measuring constructs, not just asking questions
Many things researchers study—trust, engagement, satisfaction, perceived quality—are latent constructs that cannot be captured by a single question. These require multi-item scales: several questions that together measure the underlying concept. Wherever possible, use established, previously validated scales rather than inventing your own, because validated scales come with evidence that they actually work and let your results be compared to prior research.
When you must develop a new scale, it needs proper validation—generating items, testing them, and assessing their reliability and validity before relying on them. This is the domain of psychometrics, and it is not an optional refinement: an unvalidated home-made scale is one of the fastest routes to a reviewer’s rejection. Measurement comes first, and it is worth the effort.
Sampling: who you ask decides what you can conclude
Even a perfectly worded survey produces misleading conclusions if it reaches the wrong people. The goal of sampling is a representative sample—one whose respondents reflect the population you want to draw conclusions about. When the sample is not representative, the results describe only the people who happened to answer, not the population, and generalising from them is unjustified.
Two sampling problems deserve particular attention. The first is sampling method: probability-based sampling, where members of the population have a known chance of selection, supports stronger inference than convenience sampling, where you survey whoever is easiest to reach. Convenience samples are sometimes unavoidable, but their limits must be acknowledged rather than hidden. The second is non-response bias: if the people who respond differ systematically from those who don’t—more satisfied customers more likely to answer a satisfaction survey, say—the results are skewed even with a large sample. A high response rate, and checking whether respondents differ from non-respondents, both help.
Sample size matters as well, but not in the way many assume: a large sample does not fix bias—it just gives you a precisely estimated wrong answer. Adequate size, determined by what you need to detect, matters alongside representativeness, not instead of it.
The biases to design against
Beyond question wording and sampling, several response biases can distort survey data, and good design anticipates them. Social-desirability bias leads people to answer in ways that make them look good, especially on sensitive topics—mitigated by anonymity and careful wording. Acquiescence bias is the tendency to agree with statements regardless of content, countered by mixing positively and negatively worded items. Common-method bias can inflate relationships when the same respondent provides both the predictor and the outcome in the same survey at the same time—a serious concern in much management research, addressed through design choices such as separating measurement in time or source. Naming these threats in advance lets you design around them rather than discovering them in review.
Design first, analyse later
The thread running through all of this is that survey quality is decided at the design stage, not the analysis stage. Clear questions, validated measures, a representative sample, and defences against known biases together determine whether your data can bear the weight of your conclusions. Analysis—however advanced—can only work with the data the design produced; it cannot manufacture validity that was never built in. Time spent designing the survey well is the highest-return investment in the whole project, because every later stage depends on it.
Get the design right and your survey yields data worth analysing, conclusions worth defending, and a study that stands up to review. Rush it, and the most sophisticated analysis in the world is polishing numbers that never meant what you hoped they did.
Frequently asked questions
Designing a survey or primary study?
From questionnaire design and scale validation to sampling and bias control, our team can help you collect data that is genuinely worth analysing.