Skip to content
Seven minutes. A new way to see yourself.Get your free results →
Wellington

Glossary

What is test-retest reliability, and why do scores move?

Test-retest reliability is whether the same person gets roughly the same score on a retake weeks later, not whether the statements behind one trait agree with each other in a single sitting, which is internal consistency instead. A continuous score should move only a little on a retake; a type sorted from that same score can flip entirely if it sits near the line the test cuts at.

Last updated September 16, 2026.

What is test-retest reliability, and how is it different from internal consistency?

Test-retest reliability asks a question about time: if the same person sits the same test twice, weeks or months apart, do the two scores agree? It is usually reported as a correlation between the first sitting and the second, across many people. A correlation near 1.0 means people kept almost the same rank from one sitting to the next; near zero means the two occasions have little to do with each other.

That is separate from internal consistency, which stays inside a single sitting and asks whether the statements meant to measure one trait move together, most commonly summarised as Cronbach's alpha; see what is Cronbach's alpha. A trait can be internally consistent on the day and still show a modest test-retest correlation if it is sensitive to mood, since the two numbers describe different kinds of consistency (Nunnally & Bernstein, 1994).

Why does a continuous score move a little between sittings, and why is that expected?

Classical measurement theory treats any single score as a true score plus error, and no test removes the error entirely (Nunnally & Bernstein, 1994). Some of that error is plain noise: an ambiguously worded statement, a moment of second-guessing, an ordinary night's sleep. Some is mood, since a trait like emotional stability can read differently after a stressful week than a calm one, without either reading being wrong. And some is simply reading a statement honestly a second time and noticing it means something slightly different than you assumed on the first pass. None of that is a flaw in you, and a small movement is what a working measure should produce.

How much a score wobbles also depends on how the trait was measured. A score averaged from several statements cancels out much of that noise, while a score built from only one or two statements carries far more of it forward. Using data from working adults and college students, one evaluation found very abbreviated measures, single-item ones especially, led researchers to underestimate how much traits predict and to overstate whatever else was measured alongside them; slightly longer scales improved matters substantially for little extra time (Credé et al., 2012). On a very short scale, weight the broad direction of a retake, Quiet, Balanced or High, more than the exact number.

Why can a type result flip when the underlying trait barely moved?

A trait like extraversion is continuous, and across a large sample its scores form something close to a bell curve, most people closer to the middle than either extreme. There is no natural cut-point anywhere in that distribution. A type test cuts the continuous score into two boxes at a midpoint anyway, exactly where scores are most crowded. Reinterpreting the Myers-Briggs Type Indicator against the five-factor model, one study of 468 adults aged 19 to 93 found its four indices behave as continuous dimensions relating to four of the Big Five traits, openness, extraversion, agreeableness and conscientiousness, none corresponding to emotional stability, with no support for the underlying preferences being truly dichotomous (McCrae & Costa, 1989). Someone a hair's breadth from the midpoint can be pushed across it by an ordinary wobble in one or two answers, changing their letter even though their percentile moved only slightly, a consequence of cutting a continuous scale at its most crowded point, not evidence the person changed.

A critical review of the psychometric evidence concluded that considerable caution is warranted with the four-letter formula, which does not support the inferences commonly made from it (Pittenger, 2005). A percentile does not have this failure mode, since a small wobble there produces a small change in the number, not a flip to a different category.

How stable are traits across years, measured as rank order?

A retake days or weeks later mostly tests measurement noise. Over years instead of weeks, the question becomes whether people keep their standing relative to everyone else, rank-order consistency. A review of 152 longitudinal studies, 3,217 test-retest correlations, with the gap between sittings held near 6.7 years, found this consistency rises through life: about .31 in childhood, .54 in the college years, .64 around age 30, then a plateau near .74 between ages 50 and 70 (Roberts & DelVecchio, 2000). None of those figures reaches 1.0, so adults keep much of their relative position as the decades pass, not all of it, and this population-level figure is not a claim that any individual's score cannot move; see mean-level vs rank-order change for how the group's average and an individual's position are kept as separate questions.

Is a score moving the same thing as you changing?

Not usually, and the gap is mostly a matter of scale. A retake close in time, days or a few weeks, mainly reflects measurement error: wording, mood, an honestly different reading of the same statement. A retake across months or years can also reflect real change, and researchers separate the two by asking whether the group's average moved at all, and separately whether people kept their position relative to each other; a group can shift on average while the order among its members barely changes (see mean-level vs rank-order change for both sides, with their own figures). A small movement close in time is the noise a working test should produce; a larger movement that repeats, or lines up with something that has genuinely changed, is worth reading as a signal instead.

What can make a score move between sittings, and what should you do about it?

Common reasons a personality score moves between sittings, and what each one calls for
ReasonWhat it looks likeWhat to do
Ordinary measurement errorA percentile shifts a few points with nothing different in your lifeExpect it near a band boundary; do not read one wobble as a finding
Few statements behind the traitBigger swings on a trait measured with only a couple of statementsWeight the broad band, Quiet, Balanced or High, over the exact number
Mood or context on the dayA trait tied to your current state reads differently after a hard weekRetake when nothing unusual is going on, or note the context
Re-reading a statement differentlyA statement means something slightly different the second time roundAnswer it fresh rather than recalling your last answer
Sitting near a type cut-lineA type label flips though the underlying percentile barely changedLook at the percentile behind the label, not just the letter

Reading a retake well means checking the percentile before the label and giving more weight to a change that repeats than to one surprising sitting. How to read your scores covers a percentile and a band in full, and are personality tests reliable covers reliability more broadly.

How does Wellington handle retakes?

Wellington is a science-based personality assessment from Therabot Labs LLC. It measures personality as continuous traits, built on the Big Five and HEXACO models, and writes the results back as a personal report rather than a type.

Concretely: the free Snapshot asks 60 statements, several per dimension, so one ambiguous answer does not carry a whole trait score alone. Every trait is reported as a percentile against the reference sample. The three bands (Quiet below the 30th percentile, Balanced to the 70th, High above) and the Portrait names exist to make the result easier to talk about; the percentile is the measurement, and it is always shown alongside. Norms are provisional while the norming sample is completed. A retake is a fresh sitting, not an edit to earlier answers: the Snapshot can be retaken from your dashboard at any time, the newest completed sitting drives your report, and older sittings are kept so a later one can be read against them. The Portrait is retaken as part of the Wellington Membership's annual retake, or on request.

Wellington is a wellness tool for reflection, not a medical or psychological diagnosis, and it does not replace care from a qualified professional. It is not designed for hiring or selection, and you can export or delete your data at any time.

Questions people ask

Is a personality test unreliable if my score changes on a retake?
Not necessarily. A continuous score is expected to move a little between sittings because of measurement error, wording, mood and how few statements sit behind a short scale. A large, repeated swing suggests a problem; a few percentile points either way is the kind of noise a working test produces (Nunnally & Bernstein, 1994; Credé et al., 2012).
Why did my type change but my percentile barely moved?
Because a type test cuts a continuous score into boxes at a midpoint, and that scale is continuous rather than genuinely dichotomous (McCrae & Costa, 1989). Sitting close to that line means an ordinary wobble in a couple of answers can push you across it, changing the label without much change in the score. A critical review found the formula behind such labels does not support the inferences usually drawn from it (Pittenger, 2005).
Does test-retest reliability mean personality cannot change?
No. Rank-order consistency over years is high but never perfect, rising from about .31 in childhood to a plateau near .74 between ages 50 and 70 (Roberts & DelVecchio, 2000). People keep much of their relative standing, not all of it, and that gap is exactly where real change over time shows up.

Sources

Peer-reviewed sources for the claims above. Wellington's own reliability figures will be published once the norming sample is complete.

  1. Nunnally, J. C., & Bernstein, I. H. (1994). Psychometric theory (3rd ed.). McGraw-Hill.
  2. Credé, M., Harms, P., Niehorster, S., & Gaye-Valentine, A. (2012). An evaluation of the consequences of using short measures of the Big Five personality traits. Journal of Personality and Social Psychology, 102(4), 874–888.
  3. McCrae, R. R., & Costa, P. T. (1989). Reinterpreting the Myers-Briggs Type Indicator from the perspective of the five-factor model of personality. Journal of Personality, 57(1), 17–40.
  4. Pittenger, D. J. (2005). Cautionary comments regarding the Myers-Briggs Type Indicator. Consulting Psychology Journal: Practice and Research, 57(3), 210–221.
  5. Roberts, B. W., & DelVecchio, W. F. (2000). The rank-order consistency of personality traits from childhood to old age: A quantitative review of longitudinal studies. Psychological Bulletin, 126(1), 3–25.

Read next

See your own traits, not a type.

The free Snapshot takes about seven minutes and gives you your personality card and five dimensions. No credit card, no type.

One trait a week

Not ready to take the test? Read one trait a week.

A real page from the report, one practice attached, every week. It is the easiest way to see whether Wellington reads people the way you think it should.