Do Survey Response Scales Accurately Measure Customer Intentions?

Business team reviewing a printed report with colorful bar, pie, and line charts during a collaborative office meeting.
September 1 , 2026  |  By Andy Brownback

Share this via:

Who is this research for? Marketing researchers, customer-insights teams, HR and employee-experience leaders, and other professionals who use surveys to understand or predict behavior.

Top Answer

New research suggests that survey response scales may not accurately measure customer intentions. According to these findings, people don’t interpret common survey responses such as “unlikely” and “likely” in the same way an analyst would. For example, people interpreted “unlikely” as more likely than researchers might assume, and “likely” as less likely than researchers might assume. For organizations using surveys to anticipate customer, employee, or public behavior, the findings indicate that seemingly straightforward response scales may actually introduce hidden measurement bias.

Executive Summary

Businesses routinely ask people how likely they are to buy a product, adopt a new behavior, or take some other future action. Responses such as “extremely unlikely,” “somewhat likely,” and “extremely likely” appear straightforward. But new research from Dr. Brandon R. McFadden (University of Arkansas), Emily Schlichtig (University of Arkansas alumna; Sumitomo Corporation of America), Dr. Andy Brownback (Department of Economics, Sam M. Walton College of Business), Dr. Lawton L. Nalley (University of Arkansas), and Dr. Aaron Shew (Acres) suggests that this survey design can introduce hidden measurement problems.

The researchers conducted two web-based surveys with roughly 1,200 U.S. respondents each. Participants were asked to translate familiar response categories into probabilities from 0 to 100. A common analytical assumption would place these response categories at proportionate intervals. On a five-category scale, for example, that could translate “extremely unlikely” to 10%, “somewhat unlikely” to 30%, neutral to 50%, “somewhat likely” to 70%, and “extremely likely” to 90%.

Respondents did not interpret the words that way. They generally assigned higher probabilities to unlikely categories and lower probabilities to likely categories than proportional spacing would predict. This means someone selecting an “unlikely” response may be more inclined toward the behavior than an analyst assumes, while someone selecting “likely” may be less inclined than assumed.

For decision-makers, the larger lesson is about interpretation. A survey result may look precise without being as precise as it appears. Organizations using stated intentions to inform forecasts, programs, policies, or strategy should consider whether the language in their response scales captures distinctions that respondents actually make.

Expert Insights: What should leaders know about measuring intentions?

When should leaders be most cautious about making business decisions from stated intentions?

 Dr. Brandon R. McFadden explains: “Be especially careful when a decision hinges on the difference between neighboring response options. In our research, some neighboring response options--especially on the seven-point scale--corresponded to very similar probabilities. Additionally, we observed greater variability among positive responses, like ‘extremely likely.’ So even the responses that seem most reliable may carry more uncertainty than the label suggests.”

Dr. Andy Brownback adds: “One situation where leaders should be especially cautious is when a decision depends on a shift from an ‘unlikely’ to ‘likely’ category. These seem like large shifts if we assume proportionality across categories. However, we found that respondents attach higher probabilities to ‘unlikely’ categories and lower probabilities to ‘likely’ categories than expected. That means the apparent jump in intention may be much smaller than the labels imply.”

→ Takeaway: Treat small shifts between survey categories cautiously. The labels may suggest a bigger change in intention, and more certainty, than the responses actually represent.

 How can organizations design surveys that better reflect what respondents actually mean?

Dr. Brandon R. McFadden notes: “Begin by evaluating whether each response option in a survey genuinely captures meaningful differences. We observed that respondents viewed the five categories as clearly separate; however, on a seven-category scale, some neighboring options had nearly identical probabilities. Introducing additional response choices might seem to enhance accuracy, but it could instead add unnecessary noise. If a term such as “slightly” does not establish a clear distinction in the respondent's perception, it won't produce a meaningful difference in your dataset.”

→ Takeaway: More response options do not necessarily produce better data. Choose categories that capture meaningful distinctions rather than adding choices that respondents may interpret similarly.

When might asking respondents for probabilities instead of verbal categories produce better information?

Dr. Brandon R. McFadden explains: “When the importance of a likelihood is high, requesting probabilities instead of verbal labels yields clearer information. Verbal categories are simple for respondents, but translating them into numbers relies on assumptions about their meanings--assumptions our research indicates can be incorrect. Asking for a probability on a 0-to-100 scale removes this translation, letting respondents communicate their own numerical estimate directly.

Dr. Andy Brownback adds: “By eliciting probabilities, you can capture finer distinctions than verbal categories. However, this comes at a cost. Questions take longer and require more effort from the respondents. That tradeoff may be worthwhile when distinguishing among responses in the ‘likely’ range. We found that respondents associated greater variability with these categories, suggesting that the labels convey less precise information. For example, if several analysts say it is ‘likely’ that a new competitor will enter the market, their underlying assessments may differ more than the shared label suggests. In that setting, it may be worthwhile to ask them to put numbers to their predictions.”

→ Takeaway: When precise estimates matter, consider asking for probabilities rather than relying only on verbal labels, especially for positive intentions where interpretations may vary more widely.

What should analysts consider before converting categorical survey responses into numerical scores?

Dr. Brandon R. McFadden notes: “When analysts transform ordered response categories into numerical scores, such as coding a five-point scale from 1 to 5, they often assume the response options are equally spaced. However, our data suggest that respondents might not perceive those categories as evenly spaced. Therefore, analysts should verify whether their conclusions remain consistent when using different coding methods for the scale. If the core results are significantly affected by the assumed spacing, it's crucial to recognize this before sharing the findings with decision-makers.”

→ Takeaway: Test whether results change under different ways of scoring categorical responses. If conclusions depend heavily on assumed spacing between categories, decision-makers should know that the findings are sensitive to that assumption.

Published in Journal of Behavioral and Experimental Economics (2026)

Frequently Asked Questions

Why can categorical survey scales create measurement bias?

 Categorical scales can create measurement problems when analysts treat ordered responses as though the distance between each category is equal but respondents do not interpret them that way. In this study, “unlikely” responses generally corresponded to higher probabilities than proportional spacing would predict, while “likely” responses corresponded to lower probabilities. As a result, numerical analysis based on evenly spaced categories may not perfectly represent what respondents intended to communicate. The study indicates that some apparent gaps between response categories can be partly a product of the scale itself rather than equivalent differences in people's underlying intentions.

 Are five-point survey scales better than seven-point scales?

 Not necessarily in every situation. However, this study found an important difference between the two formats. Respondents interpreted the categories on the five-category scale as distinct, while some adjacent choices on the seven-category scale were associated with similar probabilities. Respondents appeared to have particular difficulty with categories containing “slightly.” The findings suggest that adding response options does not automatically create more meaningful precision. Because the study examined specific behavioral intentions and scale wording, however, it does not establish a universal rule that organizations should always choose five categories instead of seven.

 Why are “likely” survey responses more uncertain?

 The study found that respondents generally associated greater variability with categories indicating a higher likelihood of future behavior. “Extremely unlikely” had the lowest variability, while “extremely likely” generally had the highest across the behaviors and disease-risk questions studied. The researchers suggest this pattern may reflect lower confidence when people say they are likely to change their behavior. It could therefore help explain part of the intention-action gap: a positive stated intention may contain more uncertainty than the response category alone communicates. The study does not, however, directly measure why individual respondents experienced that uncertainty.

What is the intention-action gap?

 The intention-action gap describes the difference between what people say they intend to do and what they ultimately do. Previous research cited in this study indicates that an intention to change leads to actual behavioral change only about half the time. This study offers an additional measurement-related explanation for part of that gap: researchers may sometimes misinterpret what respondents mean when they select categories such as “likely” or “unlikely.” Because positive intention categories also carried greater variability, a seemingly strong intention may contain more uncertainty than analysts recognize. The research does not claim that survey scales explain the entire intention-action gap.

Andy Brownback Andy Brownback is an associate professor of economics who specializes in behavioral and experimental economics. He uses laboratory, field, and online experiments to study questions about education and public policy. He received his bachelor's degree from Kansas State University and his PhD from the University of California, San Diego.