Skip to content
Seven minutes. A new way to see yourself.Get your free results →
Wellington

Guide

How to read your personality test scores

A personality score is a comparison, not a quantity. A percentile says what share of a reference sample scored lower than you, so the 62nd percentile means 62 percent of that sample scored lower. High is not better, it is only further from the average, and a score sitting near a band boundary can land on the other side when you take the test again.

Last updated September 16, 2026.

What do the numbers on your report actually mean?

A raw score is the average of your own answers and nothing else. If a test asks ten statements about Conscientiousness on a five-point accuracy scale and your answers average 3.8, the raw score is 3.8. On its own that says very little, because you cannot know whether 3.8 is a common answer or an unusual one.

A standardised score fixes that by putting your raw score next to a reference sample. The test subtracts the sample's average from your score and divides by the sample's spread, which gives a z score: how many standard deviations you sit from the middle. A z of zero is exactly average, and a z of plus one puts you roughly in the top sixth of people.

A percentile turns the same information into a plainer sentence. A 62nd percentile on Conscientiousness means 62 percent of the reference sample scored lower than you and about 38 percent scored higher. It does not mean you answered 62 percent of the questions correctly, because there are no correct answers, and it does not mean you are 62 percent conscientious, because the trait has no ceiling to be 62 percent of. The percentile only tells you where you stand in a line of other people.

Is a high score a good score?

Not by itself. A trait score describes where you sit, not how well you did. Some traits do relate to outcomes most people want. Conscientiousness predicts job performance across occupations (Barrick & Mount, 1991) and health-related behaviour (Bogg & Roberts, 2004), and traits as a set predict mortality, divorce and occupational attainment at a level comparable with socioeconomic status and cognitive ability (Roberts et al., 2007). Even there the finding is an average across many people rather than a verdict on one of them.

On most traits the useful direction depends on what you are doing that week. Extraversion is the strongest and most consistent trait correlate of leadership, at r = .31 across 73 samples (Judge et al., 2002), and does nothing for work that needs long stretches of solitude. Low Agreeableness is uncomfortable at a family dinner and valuable in a negotiation. If a report frames every high score as a strength and every low score as a weakness, that is a fact about its writing rather than about you.

What do the score bands mean?

Most reports group percentiles into bands, because a band is easier to talk about and the exact number carries more precision than the test can honestly support. The percentile is still the measurement, and a good report shows it alongside the band. The names matter less than the ranges, and a good report tells you the ranges.

Percentile ranges, band names and what each one is telling you
Percentile rangeCommon labelWellington bandHow to read it
1st to 15thVery lowQuietClearly below most of the reference sample. The low-pole description is the one that should sound like you.
16th to 29thLowQuietBelow average without being extreme. Expect the low-pole description to fit in some settings, not all.
30th to 70thAverageBalancedThe middle two fifths of people. You show both sides of the trait, depending on the situation.
71st to 84thHighHighAbove most of the sample. The high-pole description should be recognisable to people who know you.
85th to 99thVery highHighClearly above most of the sample. This is where a trait becomes visible to others quickly.

Balanced does not mean unremarkable, and Quiet does not mean lacking. A middle score usually means the trait moves with the situation: sociable at work and quiet at home, organised about money and loose about your inbox. That is what most people get on most dimensions.

Why does a score move when you retake the test?

Every score contains some error. Test theory treats an observed score as a true score plus noise from wording, mood, attention and which particular statements happened to be asked (Nunnally & Bernstein, 1994). A good test shrinks that noise. No test removes it.

What shrinks it most is asking more statements per trait. Reliability rises with the number of items, which is why very short scales are noisier: a measure with two items per trait buys speed at a real cost in precision, and its authors say so plainly (Gosling et al., 2003). Very abbreviated measures lead researchers to understate how much traits matter, and slightly longer scales are substantially more valid for very little extra time (Credé et al., 2012).

This is why a percentile near a band boundary is exactly the score you should expect to move. If you came out at the 31st percentile on Agreeableness in March and the 28th in June, nothing about you has changed: the measurement wobbled by three points and a boundary happened to be in the way. A change worth taking seriously is a large one, on a trait measured with several statements, that shows up more than once.

Who are you being compared with?

A percentile is meaningless until you know who is standing in the line. The same raw score can land at the 40th percentile against one sample and the 60th against another, and neither number is wrong.

The comparison matters because real differences exist between groups. Average trait levels shift across adulthood: Conscientiousness and Agreeableness both tend to rise as people get older (Srivastava et al., 2003). Averages also differ between world regions: in a survey of 17,837 people across 56 nations, respondents in South America and East Asia differed from other regions in Openness, though the authors caution against reading national league tables into the data (Schmitt et al., 2007).

So look for two things in any report: which reference sample it uses, and whether it says so in plain words. A test that norms against everyone who has ever taken it online is using a sample shaped by wherever its advertising landed. That is not disqualifying, and it is worth knowing before reading much into a handful of percentile points.

How do facets relate to the broad traits?

Broad traits are built from narrower ones called facets. Conscientiousness contains orderliness, self-discipline, achievement striving and dutifulness. The BFI-2 measures three facets beneath each of the five domains (Soto & John, 2017); the NEO inventories measure six beneath each, thirty in all (Costa & McCrae, 1992).

Facets are where an unremarkable domain score often becomes interesting. A Balanced band on Conscientiousness can be built two ways: everything near the middle, or very high achievement striving sitting beside very low orderliness. Those are different people with the same domain score. Read the domain first for the headline, then look for the facets sitting furthest from it, because those are the parts the broad score has averaged away.

Facets rest on fewer statements than the domains above them, so each carries more error. Treat a single surprising facet as a question to sit with rather than a finding, and give more weight to two or three pointing the same way.

How can you tell a real reading from a flattering one?

Bertram Forer handed his students what he told them were personality sketches drawn up for each of them individually. Every student received the same text, assembled from an astrology book, and they rated it as a highly accurate description of themselves, about 4.26 out of 5 as the demonstration is commonly reported (Forer, 1949). The sentences worked because they were true of almost everyone: you have a great deal of unused capacity, at times you are outgoing and at times reserved.

A reading doing real work behaves differently. It places you against other people rather than only describing you. It says something that does not apply to most of the population, and therefore risks being wrong. It names the costs of a trait alongside its uses. And it changes when your answers change.

  • Read the description written for the band next to yours. If it fits just as well, either you are genuinely near a boundary or the writing is loose enough to fit anyone.
  • Look for a sentence you could disagree with. Universal descriptions never contain one.
  • Check whether the report states its reference sample, its band cut points and how many questions produced each score.

How does Wellington report your scores?

Wellington is a science-based personality assessment from Therabot Labs LLC. It measures personality as continuous traits, built on the Big Five and HEXACO models, and writes the results back as a personal report rather than a type.

The free Snapshot asks 60 statements across six dimensions, Extraversion, Agreeableness, Conscientiousness, Emotional Stability, Openness to Experience, Honesty-Humility, each rated from "Very inaccurate" to "Very accurate". Each dimension is the average of its own statements, with reverse-keyed statements flipped. That average is placed against a reference sample of adults as a standard score and reported as a percentile, in one of three bands: Quiet below the 30th percentile, Balanced from the 30th to the 70th, High above the 70th. The bands and the Portrait names exist to make the result easier to talk about; the percentile is the measurement, and it is always shown alongside. The first ten questions are answered before anything is asked of you; after question ten you sign in with an email so your answers are saved and your results can be shown. There is no card. Our norms are provisional while the norming sample is completed, and the report's methods section says so. The methodology page has the full scoring description.

In the written report, every trait has three pages, one per band, written in advance by the writing and psychology team, and your answers choose which you read. The Wellington Membership extends the same scoring to 90 traits, including every Big Five facet, so you can read facets against the domain above them. The methodology shows the whole shape of it. Wellington is a wellness tool for reflection, not a medical or psychological diagnosis, and it does not replace care from a qualified professional. It is not designed for hiring, and you can export or delete your data at any time.

Questions people ask

What is a good score on the Big Five?
There is no good score, because the questions have no correct answers. A percentile tells you where you sit relative to a reference sample, not how well you did. Higher Conscientiousness relates to some outcomes many people want, but every trait carries costs as well as uses, and the most common result on most dimensions is a middle score that shifts with the situation.
Why did my results change when I retook the test?
Partly real change and mostly measurement noise. Every score is a true score plus error from wording, mood and attention, and shorter scales carry more of it. A few percentile points either way is normal, especially near a band boundary. Treat a shift as meaningful when it is large, on a trait measured with several statements, and still there the next time you look.
Does a low score mean something is wrong with me?
No. A low band describes one end of an ordinary human dimension, not a deficiency. Low Extraversion describes a person who finds long social stretches tiring and deep attention easy. Low Conscientiousness describes flexibility that costs follow-through. Each end has real uses and real costs, and a report worth reading will say what both of them are.

Sources

Peer-reviewed sources for the claims above. Wellington's own reliability figures will be published once the norming sample is complete.

  1. Nunnally, J. C., & Bernstein, I. H. (1994). Psychometric theory (3rd ed.). McGraw-Hill.
  2. Credé, M., Harms, P., Niehorster, S., & Gaye-Valentine, A. (2012). An evaluation of the consequences of using short measures of the Big Five personality traits. Journal of Personality and Social Psychology, 102(4), 874–888.
  3. Gosling, S. D., Rentfrow, P. J., & Swann, W. B. (2003). A very brief measure of the Big-Five personality domains. Journal of Research in Personality, 37(6), 504–528.
  4. Srivastava, S., John, O. P., Gosling, S. D., & Potter, J. (2003). Development of personality in early and middle adulthood: Set like plaster or persistent change? Journal of Personality and Social Psychology, 84(5), 1041–1053.
  5. Schmitt, D. P., Allik, J., McCrae, R. R., & Benet-Martínez, V. (2007). The geographic distribution of Big Five personality traits: Patterns and profiles of human self-description across 56 nations. Journal of Cross-Cultural Psychology, 38(2), 173–212.
  6. Forer, B. R. (1949). The fallacy of personal validation: A classroom demonstration of gullibility. Journal of Abnormal and Social Psychology, 44(1), 118–123.
  7. Soto, C. J., & John, O. P. (2017). The next Big Five Inventory (BFI-2): Developing and assessing a hierarchical model with 15 facets to enhance bandwidth, fidelity, and predictive power. Journal of Personality and Social Psychology, 113(1), 117–143.
  8. Costa, P. T., & McCrae, R. R. (1992). Revised NEO Personality Inventory (NEO PI-R) and NEO Five-Factor Inventory (NEO-FFI) professional manual. Psychological Assessment Resources.
  9. Barrick, M. R., & Mount, M. K. (1991). The Big Five personality dimensions and job performance: A meta-analysis. Personnel Psychology, 44(1), 1–26.
  10. Bogg, T., & Roberts, B. W. (2004). Conscientiousness and health-related behaviors: A meta-analysis of the leading behavioral contributors to mortality. Psychological Bulletin, 130(6), 887–919.
  11. Judge, T. A., Bono, J. E., Ilies, R., & Gerhardt, M. W. (2002). Personality and leadership: A qualitative and quantitative review. Journal of Applied Psychology, 87(4), 765–780.
  12. Roberts, B. W., Kuncel, N. R., Shiner, R., Caspi, A., & Goldberg, L. R. (2007). The power of personality: The comparative validity of personality traits, socioeconomic status, and cognitive ability for predicting important life outcomes. Perspectives on Psychological Science, 2(4), 313–345.

Read next

See your own traits, not a type.

The free Snapshot takes about seven minutes and gives you your personality card and five dimensions. No credit card, no type.

One trait a week

Not ready to take the test? Read one trait a week.

A real page from the report, one practice attached, every week. It is the easiest way to see whether Wellington reads people the way you think it should.