Skip to content
Seven minutes. A new way to see yourself.Get your free results →
Wellington

Glossary

What is a norm group, and why does it change your score?

A norm group, also called a norming sample or reference sample, is the group a test compares your raw score against to produce a percentile. The same answers can land at a different percentile against a different norm group, because age, gender and national averages differ on several traits, which is why an honest test names its norm group and says when its norms are still provisional.

Last updated September 16, 2026.

What is a norm group?

A norm group is the set of people a test's scoring is built from. When you answer statements about, say, Conscientiousness, the average of your ratings is a raw score, and a raw score by itself says almost nothing. It becomes informative once placed next to other people's raw scores on the same statements, and that other group is the norm group, sometimes called a norming sample or a reference sample.

Classical test theory treats this as the ordinary way a measurement gets its meaning: you do not read a score in isolation, you read it against the distribution of scores in a defined group (Nunnally & Bernstein, 1994). A norm group is not the item bank or the scoring formula; two tests can ask nearly identical questions and still report different percentiles for the same person, because their norm groups differ.

How does a norm group turn a raw score into a percentile?

The norm group supplies two numbers for each trait: an average and a spread. Your raw score is compared with that average and expressed in units of that spread, giving a standardised score. A percentile restates the same comparison in plainer language: the share of the norm group that scored lower than you.

A 65th percentile on Openness means 65 percent of the norm group scored lower than you, nothing more. It does not mean you are 65 percent open, and it says nothing about whether that is a lot in some absolute sense. The number is entirely relative to the group behind it, and that group has to be reasonably large before the comparison is stable rather than noisy.

Why can the same answers give different percentiles against different norm groups?

Because the trait averages that make up a norm group are not the same everywhere. Age is one source: Conscientiousness and Agreeableness tend to rise across early and middle adulthood, and Neuroticism declines with age for women but shows little change for men (Srivastava et al., 2003). A raw score that lands in the middle of a norm group of people in their twenties can land higher against one that also includes people in their fifties, because the older group is, on average, somewhat more conscientious and agreeable.

Gender is another source: observer ratings across 50 cultures replicated known sex differences, largest in Western cultures (McCrae & Terracciano, 2005), so a mixed group reads a score differently than one split by gender. Nationality matters too, more narrowly than popular writing suggests: a survey of 17,837 people across 56 nations found people in South America and East Asia differed from other regions specifically in Openness (Schmitt et al., 2007). None of this means your answers are wrong in one context and right in another; the comparison changed, not the trait. Does personality differ by country goes further into how fragile national averages are.

What does 'provisional norms' mean, and why does an honest test say so?

Building a proper norm group takes time: collecting raw scores from enough people, ideally described by age, gender and other relevant groups, before publishing the averages and spreads every later respondent's percentile is measured against. A test that has only just launched has not had time to do that, so provisional norms means the group is still being built and today's percentiles are a current best estimate rather than a finished, documented reference.

Saying so plainly is the honest option. Reporting a precise-looking percentile while quietly building the sample behind it makes a number look more settled than it is. A test that names its norm group and says when the norms are provisional lets you weigh the number correctly, rather than asking you to trust it on faith.

Which norm group should a test use?

There is no single right answer; every choice trades one kind of clarity for another. A broad, mixed norm group is simple and easy to explain, but averages away real differences in age, gender or region. A narrower, split norm group reads your score against people more like you, but needs a larger sample before each split is stable.

Common norm group choices, and what each one costs you
Norm group choiceWhat it tells youThe trade-off
General adult population, mixedHow you compare with adults broadly, regardless of age, gender or countryAverages away age, gender and national differences (Srivastava et al., 2003; McCrae & Terracciano, 2005; Schmitt et al., 2007)
Split by age bandHow you compare with people roughly your own ageNeeds a large sample within every band, or each one's average stays noisy
Split by genderHow you compare with your own gender's averageRoughly doubles the sample needed, and says nothing about the spread within a gender
Split by country or regionHow you compare with people from a similar cultural contextNeeds its own sample per country; most tests only manage this for a handful of nations
Whoever has taken the test so farThe fastest norm group to build, while a proper one is assembledShaped by wherever early users came from, not a deliberately built group, which is what provisional norms are

Most tests settle on a mixed general-population norm group at first, then add splits as it grows large enough to support them. The methods page sets out how scoring and bands work from there, and what a percentile means goes further into reading the number itself.

How does Wellington measure this?

Wellington is a science-based personality assessment from Therabot Labs LLC. It measures personality as continuous traits, built on the Big Five and HEXACO models, and writes the results back as a personal report rather than a type.

The free Snapshot asks 60 questions across six dimensions, Extraversion, Agreeableness, Conscientiousness, Emotional Stability, Openness to Experience, Honesty-Humility. Each dimension is the average of its own statements, placed against a reference sample as a percentile. Every trait is reported as a percentile against the reference sample. The three bands (Quiet below the 30th percentile, Balanced to the 70th, High above) and the Portrait names exist to make the result easier to talk about; the percentile is the measurement, and it is always shown alongside. Our norms are provisional until the norming sample is complete, and we say so rather than presenting an early percentile as a finished one; the methods page describes the current position.

The Wellington Membership extends the same scoring to 90 traits, including every Big Five facet, and continues toward all 255 traits across eight domains. Wellington is a wellness tool for reflection, not a medical or psychological diagnosis, and it does not replace care from a qualified professional. It is not designed for hiring, and you can export or delete your data at any time.

Questions people ask

Is a norm group the same thing as a sample size?
No. A sample size is just how many people were measured. A norm group is that group used specifically as the comparison for scoring, described by who they are, not only by how many there were. A large sample skewed heavily toward one age or country is still a narrow norm group, whatever the total number.
Why did my percentile change when I retook the test?
Two reasons, usually together. Your raw score can move a little from ordinary variation in mood or attention, and if the norm group is still being built, its average and spread can shift slightly as more people are added. A small change is expected; a large one is worth attention.
Does a bigger norm group always give a more accurate percentile?
Generally, up to a point. A bigger norm group gives a steadier average and spread, so the percentile it produces wobbles less over time. Beyond a certain size the gains get smaller, and who is in the group matters as much as how many people are in it.
How can I tell what norm group a test is using?
A test that takes this seriously will usually describe its reference sample somewhere, in a methods page or a report footer, including roughly who is in it and whether the norms are considered finished or provisional. If a test never mentions its comparison group, treat any percentile it reports with some caution.

Sources

Peer-reviewed sources for the claims above. Wellington's own reliability figures will be published once the norming sample is complete.

  1. Nunnally, J. C., & Bernstein, I. H. (1994). Psychometric theory (3rd ed.). McGraw-Hill.
  2. Srivastava, S., John, O. P., Gosling, S. D., & Potter, J. (2003). Development of personality in early and middle adulthood: Set like plaster or persistent change? Journal of Personality and Social Psychology, 84(5), 1041–1053.
  3. Schmitt, D. P., Allik, J., McCrae, R. R., & Benet-Martínez, V. (2007). The geographic distribution of Big Five personality traits: Patterns and profiles of human self-description across 56 nations. Journal of Cross-Cultural Psychology, 38(2), 173–212.
  4. McCrae, R. R., & Terracciano, A. (2005). Universal features of personality traits from the observer's perspective: Data from 50 cultures. Journal of Personality and Social Psychology, 88(3), 547–561.

Read next

See your own traits, not a type.

The free Snapshot takes about seven minutes and gives you your personality card and five dimensions. No credit card, no type.

One trait a week

Not ready to take the test? Read one trait a week.

A real page from the report, one practice attached, every week. It is the easiest way to see whether Wellington reads people the way you think it should.