Forms · Glossary

What is a rating scale question?

A rating scale question asks somebody to judge one thing on an ordered set of points, from low to high. The points may be numbers, stars, or words such as poor through to excellent. It differs from agreement scales because the respondent is grading the object directly rather than saying how much they agree with a claim about it.

Rating scales are the most used question type in feedback of any kind, and the two decisions that shape the data are made before a single response arrives: how many points, and what the points are called.

· Co-founder

6 min read · Published

How many points, and what each suits
PointsSuitsTrade off
3Children, very short surveys, low literacy audiencesToo coarse to show movement over time
5Most feedback and satisfaction questionsThe default for good reasons, but compresses strong opinions
7Research where fine distinctions matterHarder to label in words, more effort per item
10 or 11Benchmarked metrics and international comparisonPeople use the scale very differently across cultures
StarsConsumer contexts where the convention is familiarSkews high, and the meaning of each star is unstated
Even numbersForcing a sideRemoves the honest middle as well as the lazy one

Labelling every point beats labelling the ends

A scale with only the two ends labelled leaves every point between them to interpretation, and different people interpret them differently. Labelling all the points in words, such as poor, fair, good, very good and excellent, gives everybody the same reference and makes the results comparable between respondents and across waves. The main objection is that word labels are harder to pick for seven or ten points, which is a real constraint and one of the better arguments for five. Keep the labels evenly spaced in meaning, which is harder than it sounds: poor, acceptable, good, excellent has a gap at the bottom and a crowd at the top. Avoid mixing numbers and words in a way that implies more precision than the question can carry.

What a mean rating hides

Rating data is ordinal, which means the distance between good and very good is not guaranteed to equal the distance between fair and good. Averages are calculated anyway, everywhere, and they are useful for tracking, but they should not be the only number reported. A mean of four point one can be everybody at four, or a split between fives and twos, and those are different situations requiring different responses. Report the distribution alongside the mean, or report the share who chose the top two points, which is easier for a non technical audience to interpret and moves for reasons that can be explained. Be careful comparing means across questions with different numbers of points, which is a common error in dashboards.

Why ratings drift upward

Ratings collected from people who chose to respond skew high, because the moderately satisfied majority rarely bothers. Add the tendency to be polite towards a named business, and a five point scale often produces a distribution crammed into the top two points, which leaves nothing to measure. Three things help. Ask about a specific interaction rather than the relationship in general, since specificity produces sharper answers. Use word labels that make the middle sound acceptable rather than failing. And track the share choosing the top point rather than the mean, because that number still moves when the average has flattened against the ceiling.

Keeping scales consistent across a form

Changing scale direction or length within one form is a reliable way to collect confused answers. If one question runs poor to excellent and the next runs excellent to poor, some people will answer the second on the first question's pattern, and there is no way to detect it afterwards. Pick one direction and one length for the whole form and hold to it. The same applies across waves of a recurring survey: changing a five point scale to a seven point one breaks the comparison with every earlier round, and no amount of rescaling afterwards fully repairs it. If a change is genuinely needed, run both versions once in parallel so the two series can be bridged.

Building it in a form

A single rating question is a radio field with the scale points as options, laid out in one, two or three columns depending on how many points there are and how long the labels run. Where several things share one scale, a matrix is the compact build: rows by columns, with radio style allowing one answer per row. A matrix exports as one column with each statement and its answer joined as Statement: Answer, so a rating you plan to chart on its own is better asked as its own field as well. Radio fields export the option labels rather than the stored values.

Deciding what the scale is attached to

Before choosing points and labels, settle what exactly is being rated, because a scale attached to a vague object collects vague answers. How would you rate our service invites the respondent to average everything they can remember, and two people answering four will have averaged different things. How would you rate the time it took to get a reply produces an answer both of them mean the same way. The same discipline applies to time frame: rating an experience from six months ago and one from yesterday on the same scale puts two different kinds of memory in one column. Name the period in the question. Where several aspects need rating, list them as separate short items rather than folding them into one general question, and accept that four specific ratings are more useful than one broad one even though they take longer to ask.

Questions people ask

Five points or seven?

Five for general feedback, seven where you need finer distinctions and the audience is engaged. Seven gives a little more room to move and a little more noise, and it is harder to label in words. For a public facing satisfaction question, five labelled points will serve almost every purpose and will be finished by more people.

Should the scale start at the positive end?

Convention in most western surveys runs from negative on the left to positive on the right, matching how people read a number line. What matters most is consistency within the form and across waves. If you flip the direction between rounds, the comparison is broken and the break will not be visible in the numbers.

Are star ratings a rating scale?

They are, with the labels removed. Stars are familiar and quick, and they skew high because the cultural meaning of four stars is closer to good than to average. They also give you no information about what each star meant to the person. Use them where the convention is expected and word labels where the data matters.

Can ratings be averaged across different questions?

Only if the questions share the same scale and are genuinely measuring related things. Averaging satisfaction with delivery and satisfaction with the product gives a number that changes for two reasons at once. If a summary index is needed, build it deliberately and keep reporting the components alongside it.

Should a rating question be required?

On a short feedback form, usually yes, since an unanswered rating is the whole point of the question. On a longer survey, requiring every rating pushes people who have no view into giving one, which adds noise. Adding a not applicable option is the usual compromise where an item may not apply to everybody.

Do numbers without labels work?

They work when the scale is very familiar, such as a nought to ten recommendation question, and they work badly otherwise. Without labels, one person's seven is another person's five, and the answers cannot be compared across respondents. If you use numbers, at minimum label both ends and keep the direction stable.

Make one with forms

The button opens the generator with this use case already described. Change the wording to match your own.

Create a form with OneCraft

Related questions

Step by step in the builder: Every form field and when to use it, then Create a form from scratch with AI.

Sources

Written and checked by the OneCraft team. Last checked .