Forms · Glossary

What is a representative sample?

A representative sample is a group of respondents whose makeup matches the wider population on the characteristics that affect the answers being measured, such as age, location or type of customer. It is what lets a result from a few hundred people describe thousands. Size alone does not make a sample representative; how people were selected does.

Most survey mistakes that survive into decisions are sampling mistakes rather than wording mistakes. A carefully written questionnaire sent to the wrong slice of people still produces a confident, wrong answer.

· Co-founder

5 min read · Published

Common ways to draw a sample
MethodHow people are chosenLikely to be representativeTypical risk
Simple randomEvery member of a complete list has an equal chance of selectionYes, if the list is complete and people respondNon-response undoing the randomness
Stratified randomThe population is split into groups and people are picked at random within eachYes, and small groups are guaranteed a placeNeeds the group of every person known in advance
SystematicA random starting point, then every nth person on a listUsuallyA list ordered in a repeating pattern skews the selection
ConvenienceWhoever is easy to reach, such as a link shared onlineRarelyOverrepresents the most engaged and most available people
QuotaSet numbers per group, filled without random selectionOnly on the quota traitsBias inside each group stays hidden
SnowballRespondents recruit people they knowNoNetworks cluster people who are alike

Representative on which traits

No sample matches a population on everything, so the useful question is which traits drive the answers you are measuring. A survey about public transport needs the right mix of people who live near train lines; a survey about a software product needs the right mix of heavy and occasional users. Matching on age and gender while missing the trait that matters produces a sample that looks balanced in the methods section and misleads in the results. Pew Research Center showed how large the gap can be. In 2016 it compared nine online samples from eight vendors against 20 benchmarks from high quality government sources. Average estimated bias on those benchmarks was 15.1 points for Hispanic adults and 11.3 points for Black adults, and every sample overrepresented politically and civically engaged people.

Probability and non-probability samples

Sampling methods split into two families. In a probability sample, as Qualtrics's guide puts it, each member of the population has a known, non-zero chance of being selected, which is what allows a margin of error to be calculated and a result to be generalised. Simple random, stratified, systematic and cluster sampling all belong here. In a non-probability sample the researcher, or the respondents themselves, decide who takes part on grounds such as convenience or availability, and the same guide warns these are unlikely to produce a representative sample. Non-probability samples are cheaper and faster, and for early exploration they are often good enough. Quota sampling tries to bridge the gap by filling set numbers per group, which fixes the proportions you set quotas on and leaves every other difference untouched. The vendor that performed best in Pew's comparison relied on an elaborate set of adjustments, not luck.

The sampling frame comes first

Before any method applies there has to be a list, or some other route, that reaches the population. That list is the sampling frame, and a sample can only represent the people on it. A customer database that holds only people who created accounts misses everyone who paid in cash. A staff directory can miss contractors. Pew's American Trends Panel shows the effort a strong frame takes: since 2018 it has recruited by selecting households from the U.S. Postal Service's master list of residential addresses, it adds new members each year because some stop responding, and it retires members from groups that have become overrepresented. For an ordinary business survey the lesson is smaller but the same. Write down who is on the list, who is not, and whether the missing people might answer differently.

Checking a sample you already have

Once responses are in, compare the sample's makeup with known facts about the population. A staff survey can be checked against headcount by team; a customer survey against the share of customers by region or product. Where the sample is off on a trait that matters, weighting can rebalance it, provided the population figures are reliable and no group is almost absent. Where a key group barely responded, weighting stretches a handful of answers too far, and the better report shows that group's count and says its result is uncertain. It also helps to compare early and late respondents, since the people who needed reminders are often closer to the people who never answered at all.

Keeping a form inside its frame

A public form link can be forwarded anywhere, which turns a careful random selection into a convenience sample one share at a time. When the frame is a known group, the login required access type limits the form to an allowed list of addresses, or to whole domains written like @company.com. That check happens when the form opens; the account is not saved with the answers, so it keeps outsiders out without identifying anyone. Ask the traits you will compare against population figures as single dropdown or radio questions, because each becomes its own column in the CSV export.

Questions people ask

Can 200 people represent a whole country?

Yes, if they were chosen at random from a good frame and most of them responded, although results carry a margin of about 6.9 points at 95 percent confidence. What 200 people cannot do is represent small groups within that country reliably. The size limits precision; the selection method decides whether the result points in the right direction at all.

What is sampling bias?

It is a systematic difference between the sample and the population caused by how people were selected, such as surveying shoppers only on weekday mornings or recruiting through one social network. Unlike random sampling error, it does not shrink as the sample grows. Ten thousand responses from a biased route are simply a precise estimate of the wrong group.

Are online surveys ever representative?

They can be when recruitment is random and offline, as with panels that invite households chosen at random from address lists and then survey them online. They rarely are when anyone who sees a link can take part. The mode of answering matters much less than how people were invited to answer.

What is coverage error?

Coverage error is the gap between the sampling frame and the population. People who are missing from the list cannot be selected, however carefully the sample is drawn. Surveying customers by email misses customers with no address on file, and if those customers are older or buy differently, every result inherits the gap.

Should I use quotas for a small survey?

Quotas are a reasonable way to make sure small but important groups are heard, especially when a random sample would leave them with a handful of responses. Set quotas only on traits you have reliable population figures for, stop inviting a group once its quota is full, and describe the sample as quota based rather than random.

How do I describe a sample that is not representative?

Plainly. Say how people were recruited, how many answered, and how their makeup compares with the population where you know it. Present results as what these respondents said rather than what customers or staff think, and leave out the margin of error. Readers can use honest findings from a limited sample; they cannot use a misleading generalisation.

Make one with forms

The button opens the generator with this use case already described. Change the wording to match your own.

Create a form with OneCraft

Related questions

Step by step in the builder: Create a form from scratch with AI, then Control who can fill in your form.

Sources

Written and checked by the OneCraft team. Last checked .