Skip to content
Seven minutes. A new way to see yourself.Get your free results →
Wellington

Guide

Are personality tests reliable? Reliability, explained

Reliability means a test gives a consistent score: the statements meant to measure a trait move together, and a retake weeks later lands close to the first score. It is not the same as validity, which is whether the score means what it claims (Nunnally & Bernstein, 1994). A trait measured with several statements is usually reliable; a type sorted from a continuous score is not, by design.

Last updated September 16, 2026.

What does reliability mean, and how is it different from validity?

Reliability is a question about consistency, asked two ways. Internal consistency asks whether the statements meant to measure one trait move together within a single sitting. Test-retest reliability asks whether the same person gets roughly the same score on a retake weeks later, with nothing in their life to explain a change. A score that swings for no reason is not tracking anything stable (Nunnally & Bernstein, 1994).

Validity is separate: whether the score means what the label says. A bathroom scale that reads three kilograms heavy every time is reliable, consistent, but not valid, since the number does not match your weight. A personality score has to be reliable before it can be valid, but reliability alone does not prove the trait relates to anything real; that is covered in how accurate are personality tests.

Why do more questions make a test more reliable?

Every statement carries some noise: an odd mood that day, an ambiguously worded question, a moment of second-guessing. Averaging several statements about the same trait cancels out much of that noise, because one statement's quirks rarely match another's. A score built from one statement carries all of that noise; a score built from six or eight carries far less.

Using data from working adults and college students, one evaluation found that very abbreviated measures, single-item measures especially, led researchers to underestimate how much personality traits actually predict, and to overstate the importance of the newer constructs measured alongside them. Slightly longer scales meaningfully improved accuracy for little extra time from the person answering (Credé et al., 2012).

Short scales still have a place. A five- or ten-item Big Five measure, built for settings where time is scarce, reaches an adequate level on test-retest reliability and predicting known outcomes, but its own authors call it somewhat inferior to standard multi-item measures and recommend it only where the alternative is no measurement at all (Gosling, Rentfrow & Swann, 2003). The fewer statements behind a score, the more weight to put on its broad direction, Quiet, Balanced or High, and less on the exact number.

How stable are traits across years?

A related question is whether who you are now predicts who you will be years from now. The evidence concerns rank-order consistency: whether your standing relative to other people holds up over time, not whether your score never moves.

A review of 152 longitudinal studies, 3,217 test-retest correlations with the gap between sittings held near 6.7 years, found this consistency rises through life: about .31 in childhood, .54 in the college years, .64 around age 30, then a plateau near .74 between 50 and 70 (Roberts & DelVecchio, 2000). None of those figures is 1.0. Adults keep much of their relative position on a trait, not all of it; this is stability in relative standing, not evidence that personality stops changing after 30.

Why do type tests retest poorly?

This is where reliability breaks down for a whole category of test, not because the questions are bad, but because of what happens after scoring. A trait like Extraversion is continuous: across a large sample, scores form something close to a bell curve, most people near the middle, fewer toward either end.

A type test cuts that continuous score at a midpoint, exactly where the scores are most crowded, and sorts everyone into one of two boxes. Someone sitting a hair's breadth from the line can be pushed to the other side by an ordinary wobble in one or two answers, changing their letter and the whole description that comes with it. This is a structural consequence of cutting a continuous scale, not a flaw specific to one product; a critical review of the instrument most associated with this format concludes its four-letter formula does not support the inferences commonly drawn from it (Pittenger, 2005). A score reported as a percentile lacks this failure mode: a small wobble produces a small change, not a flip to a different category.

What kinds of reliability show up on a good report?

The main kinds of reliability referenced on personality reports, and what each one actually checks
Kind of reliabilityWhat it checksWhat a low number would mean
Internal consistencyWhether the statements meant to measure one trait move together in a single sittingThe statements may measure more than one thing, or the trait is defined too broadly
Test-retest reliabilityWhether the same person gets a similar score weeks or months apart, with no real change betweenThe score is noisy, or sensitive to mood or wording on the day
Rank-order stability (population level)Whether people keep their relative standing on a trait over years, across a groupIndividual differences would not be predictable from earlier scores at all
Self-observer agreementWhether a self-report lines up with what someone who knows you well would sayThe trait may be hard to see from outside, or hard to see in yourself

A single number cannot cover all four, so a report worth trusting says which kind it is describing. Once a test is reliable in this sense, the next question is what its scores mean; see how to read your scores and, for how a test was built and checked, the methodology page.

How does Wellington measure this?

Wellington is a science-based personality assessment from Therabot Labs LLC. It measures personality as continuous traits, built on the Big Five and HEXACO models, and writes the results back as a personal report rather than a type.

Concretely: the free Snapshot measures Extraversion, Agreeableness, Conscientiousness, Emotional Stability, Openness to Experience, Honesty-Humility across 60 statements, several per dimension, so one ambiguous answer does not carry a whole trait score alone. The Wellington Membership extends this to 90 traits, with enough statements per trait to measure facets rather than broad dimensions alone. Every trait is reported as a percentile against the reference sample. The three bands (Quiet below the 30th percentile, Balanced to the 70th, High above) and the Portrait names make the result easier to talk about; the percentile is the measurement, and it is always shown alongside.

We have not yet published our own reliability figures. We would rather say that plainly than quote numbers we have not earned; they will appear on the methodology page once the norming sample is complete, and the norms stay provisional until then. Wellington is a wellness tool for reflection, not a medical or psychological diagnosis, and it does not replace care from a qualified professional. It is not designed for hiring or selection, and you can export or delete your data at any time.

Questions people ask

Is a personality test reliable if I get a different result the second time?
It depends how different. A trait reported as a continuous score should move only a little on a retake taken close in time, since some day-to-day variation is normal. A large swing suggests the statements behind that trait are too few, or something genuinely changed. A type test flipping to a different letter is not necessarily a sign of change; it can just mean your score sits near the midpoint the test cuts at.
Does a longer personality test mean a more reliable one?
Generally yes, up to a point. Averaging more statements about the same trait cancels out more noise in any single answer, which is why very short measures trade precision for speed, while slightly longer scales improve accuracy substantially for little extra time (Credé et al., 2012; Gosling et al., 2003). Past a reasonable number of statements per trait, the gains taper off.
Can a reliable personality test still be wrong about me?
Yes. Reliability only means the test is consistent, not that the score is accurate. A test can give the same result every time and still be measuring something other than what it claims. Reliability is a precondition for a useful test, not proof it is valid, a separate question covered in how accurate personality tests are.
Why do MBTI-style type tests seem to change results so often?
Because the underlying scales are continuous and most people sit closer to the middle than either extreme. Cutting a continuous scale into two boxes at that midpoint means an ordinary wobble in a few answers can move someone across the line, changing their letter and their whole description, even though a review of the instrument concluded its formula does not support the inferences drawn from it (Pittenger, 2005).

Sources

Peer-reviewed sources for the claims above. Wellington's own reliability figures will be published once the norming sample is complete.

  1. Nunnally, J. C., & Bernstein, I. H. (1994). Psychometric theory (3rd ed.). McGraw-Hill.
  2. Credé, M., Harms, P., Niehorster, S., & Gaye-Valentine, A. (2012). An evaluation of the consequences of using short measures of the Big Five personality traits. Journal of Personality and Social Psychology, 102(4), 874–888.
  3. Gosling, S. D., Rentfrow, P. J., & Swann, W. B. (2003). A very brief measure of the Big-Five personality domains. Journal of Research in Personality, 37(6), 504–528.
  4. Roberts, B. W., & DelVecchio, W. F. (2000). The rank-order consistency of personality traits from childhood to old age: A quantitative review of longitudinal studies. Psychological Bulletin, 126(1), 3–25.
  5. Pittenger, D. J. (2005). Cautionary comments regarding the Myers-Briggs Type Indicator. Consulting Psychology Journal: Practice and Research, 57(3), 210–221.

Read next

See your own traits, not a type.

The free Snapshot takes about seven minutes and gives you your personality card and five dimensions. No credit card, no type.

One trait a week

Not ready to take the test? Read one trait a week.

A real page from the report, one practice attached, every week. It is the easiest way to see whether Wellington reads people the way you think it should.