Glossary
What is a Likert scale, and why do personality tests use one?
A Likert scale is a set of ordered response options, such as "Very inaccurate", "Moderately inaccurate", "Neither accurate nor inaccurate", "Moderately accurate" and "Very accurate", or the more familiar strongly disagree to strongly agree, that turns how well a statement describes you into a number. Tests use one because several such numbers can be averaged into a single trait score.
Last updated September 16, 2026.
What is a Likert scale?
A Likert scale is a set of ordered response options attached to a statement, running from one extreme to its opposite with a neutral point in the middle. It is named after the psychologist Rensis Likert, who introduced it in the 1930s as a way of turning attitudes into numbers, and it has since become the standard way personality tests ask their questions.
On a trait-based personality test, the statement is usually about behaviour rather than opinion, and the options describe how accurately it fits you rather than how much you agree with it. Wellington's own scale runs "Very inaccurate", "Moderately inaccurate", "Neither accurate nor inaccurate", "Moderately accurate" and "Very accurate". Other inventories phrase the same idea as agreement instead of accuracy, from strongly disagree to strongly agree. The wording changes; the underlying structure, an ordered set of points between two extremes, does not.
What makes it a Likert scale, rather than a simple yes or no, is that the points are ordered and roughly evenly spaced in the respondent's mind: moving from "Very inaccurate" to "Neither accurate nor inaccurate" should feel like a similar step to moving from "Neither accurate nor inaccurate" to "Very accurate". That assumption is what lets a test average several answers together and treat the result as a number rather than a category.
Why do tests use five or seven points, and what does the middle option do?
Most personality tests settle on five or seven points. Fewer than that and the scale loses information, because people who feel mildly different about a statement are forced into the same answer as people who feel strongly different. More than about seven and the extra points rarely help, because most people cannot reliably tell "quite accurate" and "very accurate" apart, so the scale gains complexity without gaining precision.
The middle option, "Neither accurate nor inaccurate" on Wellington's scale, does a specific job. Some statements genuinely sit halfway true for a given person, true in some settings and not others, and the midpoint lets them say so instead of being pushed to a side that does not fit. It also gets used, less usefully, to dodge a statement someone could actually place. A test cannot tell the two uses apart from the answer alone, which is why some tests drop the middle option and force a side instead.
| Format | Points | Trade-off |
|---|---|---|
| Five-point accuracy or agreement | 5 | The common choice. Enough range to be useful, few enough that people can place themselves quickly. |
| Seven-point | 7 | Finer distinctions at the edges, at the cost of a small amount of extra time and, for some people, more indecision. |
| Forced-choice, no midpoint | 4 or 6 | Removes the fence-sitting option, but a genuinely neutral person is forced to pick a side, which adds noise of its own. |
| Visual slider or continuous scale | Effectively unlimited | More granular in principle, but harder for people to use consistently without a labelled anchor at each end. |
How do individual answers become a trait score?
A single statement is never the whole test. Several statements tap the same trait from different angles, and some are deliberately reverse-keyed, meaning agreement points toward the low end of the trait rather than the high end. Reverse-keyed items covers why that matters and how the flip works; in short, reversed answers are turned around before scoring, so every statement measuring a trait ends up pointing the same way, and the trait score is the average of all of them.
Averaging matters because one statement on its own is a thin instrument. It carries whatever the respondent took a particular word to mean, their mood while answering, and the specific behaviour the statement happened to name, on top of the trait it was meant to capture. Classical measurement theory treats this as ordinary: reliability, the consistency of a score across repeated measurement, rises as more items tapping the same underlying trait are added and averaged (Nunnally & Bernstein, 1994).
The cost of ignoring that has been measured rather than assumed. Using data from 437 employees and 355 college students, Credé et al. (2012) found that very abbreviated Big Five measures, single-item ones worst of all, led researchers to understate how much the traits actually mattered for behaviour, while scales only slightly longer improved the picture substantially for very little extra time from the person answering. A Likert scale is what makes that averaging possible: without ordered, comparable points on each statement, there is nothing to add together.
Where do these statements come from, and what can go wrong when people answer them?
Many Likert-scale statements used across free and commercial tests trace back to the International Personality Item Pool, a public-domain collection built so broad traits could be measured without licensing a commercial inventory. Seven measurement researchers set out the case for public-domain scales, using that pool as their example (Goldberg et al., 2006). Because the items are free to copy and reuse, near-identical statements and five-point scales turn up across many tests.
Published inventories build on the same format with their own items. The BFI-2, a widely used research measure, asks 60 statements on a Likert scale, 12 per domain and 4 per facet, organised as fifteen facets under five broad domains, a hierarchy validated against both self-reports and reports from people who know the respondent (Soto & John, 2017). Its authors make it available at no cost for research and educational use. What differs between tests is which statements were kept, how many per trait, and how carefully it was checked, not the scale format itself.
A Likert scale also opens the door to response habits unrelated to the trait being measured. Acquiescence is the tendency to agree with most statements regardless of content, one reason tests mix in reverse-keyed items. Extreme responding, and its opposite, avoiding the endpoints altogether, both distort a person's true position. Neither means someone is answering in bad faith; they are known habits careful scoring tries to work around.
How do you answer a Likert-scale question well?
- Answer with your first honest reading. The response you would give in a few seconds usually reflects the habit better than one you reason your way toward.
- Use the middle option when it is genuinely true. Plenty of statements really do not fit either way, and saying so is information rather than a dodge.
- Think about your typical self, not one occasion. A single unusual week is a poor guide to where you sit on an ordinary scale point.
- Do not guess which end sounds better. The statements have no built-in verdict, so there is nothing to win by leaning toward either extreme.
- Read each statement on its own terms. Reverse-keyed items are mixed in on purpose, so matching the tone of the last answer will not serve you.
These habits matter most on longer inventories, where the same trait is asked about from several angles across the personality test questions you will see. How the whole set is built and turned into a report is covered in methodology, and what a score built this way can and cannot tell you is covered in are personality tests reliable.
How does Wellington use a Likert scale?
Wellington is a science-based personality assessment from Therabot Labs LLC. It measures personality as continuous traits, built on the Big Five and HEXACO models, and writes the results back as a personal report rather than a type.
The free Snapshot rates 60 statements on the five-point accuracy scale used above, from "Very inaccurate" to "Very accurate", across the 6 dimensions it reports (Extraversion, Agreeableness, Conscientiousness, Emotional Stability, Openness to Experience, Honesty-Humility), and takes about 7 minutes. The statements themselves are real public-domain IPIP items and public-domain items written in the same style, curated for a wellness setting, and no items are licensed from a commercial inventory. Several statements are asked per dimension, some keyed positively and some in reverse, and reversed answers are flipped before the statements for a dimension are averaged. Every trait is reported as a percentile against the reference sample. The three bands (Quiet below the 30th percentile, Balanced to the 70th, High above) and the Portrait names exist to make the result easier to talk about; the percentile is the measurement, and it is always shown alongside. Our norms are provisional while the norming sample is completed. The first ten questions are answered before anything is asked of you; after question ten you sign in with an email so your answers are saved and your results can be shown. There is no card.
The Wellington Membership adds a further 336 statements on the same scale to reach 90 traits, including every Big Five facet and all 24 character strengths. Wellington is a wellness tool for reflection, not a medical or psychological diagnosis, and it does not replace care from a qualified professional. It is not designed for hiring, and you can export or delete your answers at any time.
Questions people ask
- Is a Likert scale the same as an agree-disagree scale?
- An agree-disagree scale is one common version of a Likert scale, with endpoints of strongly disagree and strongly agree. Other versions use accuracy instead, from very inaccurate to very accurate, or frequency, from never to always. All share the same structure: ordered points between two extremes, averaged into a score.
- Why is a 5-point Likert scale so common?
- Five points balance two needs: enough range that people who feel differently can say so, and few enough that most respondents can reliably place themselves without distinguishing points that feel almost identical. Seven-point scales add a little more range at a small cost in speed, which is why both remain common.
- Does choosing the middle option mean I am undecided?
- Not necessarily. The middle option exists because some statements genuinely sit halfway true for a given person, and using it to say so is honest rather than evasive. It is a problem only when used to dodge a statement someone could actually place, which is why some shorter tests remove it instead.
- Why do some Likert-scale statements feel worded backwards?
- Those are reverse-keyed items, included on purpose so a habit of agreeing with everything shows up as an inconsistent pattern rather than a high score. Reverse-keyed items explains how scoring flips these answers before they join the average.
- Can a Likert scale be wrong or misleading?
- A Likert scale is only as good as the statements attached to it and how many measure a given trait. A single vague statement carries a lot of noise regardless of the scale, which is why careful tests use several statements per trait rather than relying on the response format alone (Credé et al., 2012).
Sources
Peer-reviewed sources for the claims above. Wellington's own reliability figures will be published once the norming sample is complete.
- Nunnally, J. C., & Bernstein, I. H. (1994). Psychometric theory (3rd ed.). McGraw-Hill.
- Credé, M., Harms, P., Niehorster, S., & Gaye-Valentine, A. (2012). An evaluation of the consequences of using short measures of the Big Five personality traits. Journal of Personality and Social Psychology, 102(4), 874–888.
- Goldberg, L. R., Johnson, J. A., Eber, H. W., Hogan, R., Ashton, M. C., Cloninger, C. R., & Gough, H. G. (2006). The international personality item pool and the future of public-domain personality measures. Journal of Research in Personality, 40(1), 84–96.
- Soto, C. J., & John, O. P. (2017). The next Big Five Inventory (BFI-2): Developing and assessing a hierarchical model with 15 facets to enhance bandwidth, fidelity, and predictive power. Journal of Personality and Social Psychology, 113(1), 117–143.
Read next
See your own traits, not a type.
The free Snapshot takes about seven minutes and gives you your personality card and five dimensions. No credit card, no type.