Skip to content
Seven minutes. A new way to see yourself.Get your free results →
Wellington

Guide

Personality test validity: what it means and how to check it

A personality test has validity when it measures what it claims to measure, not just when it feels accurate. That claim breaks into parts: do the items cover the trait, does the trait behave the way the theory predicts, do independent measures of it agree, and do scores forecast anything real. A test that can point to evidence for each part is valid; one that only feels true is not.

Last updated September 16, 2026.

What does it mean for a personality test to be valid?

Validity and reliability answer different questions. Reliability asks whether a test gives the same answer twice, on a retake or across similar items. Validity asks whether that answer means what it claims to mean. A bathroom scale that reads five pounds heavy every time is reliable and not valid. Classical measurement theory treats validity as several related questions rather than one; the standard graduate account of that framework is Nunnally and Bernstein (1994).

Four kinds of validity come up most often in personality research. Content validity asks whether the items sample the whole trait, so a scale for Conscientiousness should ask about tidiness, follow-through and planning, not just tidiness. Construct validity asks whether the trait behaves the way the underlying theory says it should, holding together as a stable pattern across samples and over time. Convergent validity asks whether independent ways of measuring the same trait agree with each other, such as a self-report and a rating from someone who knows the person well. Predictive validity asks whether a score measured today relates to something that happens later.

Four questions a valid personality test should be able to answer
Kind of validityThe question it answersWhat evidence looks like
Content validityDo the items cover the whole trait, not one slice of it?The item set spans the trait's known facets rather than repeating one theme
Construct validityDoes the trait behave the way the theory predicts?The same structure recurs across independent samples and holds up over time
Convergent validityDo different ways of measuring the trait agree?Self-reports and observer ratings, or two separate instruments, correlate substantially
Predictive validityDo scores forecast something that matters later?A score measured now relates to an outcome measured months or years later

Does the Big Five have structural validity?

For the Big Five, construct validity mostly means the same five-factor pattern keeps turning up. Goldberg (1993) traces that pattern across decades of lexical research, work that began with lists of trait words rather than a single questionnaire, and notes that the model has survived repeated attempts by critics to replace it with something else. A structure that only appears in one dataset, built by one team, is much easier to doubt than one that keeps reappearing under different methods.

Convergent validity for the Big Five has direct evidence too. McCrae and Costa (1987) compared self-reports with peer ratings, using both adjective checklists and questionnaire scales. The same five factors turned up in both self-report and observer data, and mean peer ratings correlated with self-reports at .25 to .62 across the five factors, which the authors describe as substantial cross-observer agreement. That is what convergent validity looks like in practice: different vantage points on the same person landing in roughly the same place. The full case for the Big Five's structure goes further into how that evidence built up.

Does the Big Five predict anything real?

Structural agreement is not the same as usefulness. The sharper test is predictive validity: does a score from a questionnaire relate to anything that happens afterward. Roberts et al. (2007) reviewed prospective longitudinal studies, ones that measured personality first and an outcome years later, and compared personality's predictive power against socioeconomic status and cognitive ability for three outcomes: mortality, divorce and occupational attainment. The size of personality's effect on all three was not distinguishable from the size of the SES or IQ effect, and the authors argue the effects are not trivially small.

A fair question follows: how much of that literature holds up under a stricter test. Soto (2019) ran a preregistered, high-powered replication project on 78 previously published trait-outcome associations. Eighty-seven percent replicated in the expected direction, though the replicated effects were typically about 77 percent as strong as the original ones, a reasonably accurate map of real relationships rather than one built on associations that vanish under scrutiny.

Do other people's ratings agree with a person's own?

One more kind of convergent evidence comes from comparing a person's self-report with what people who know them say. Connelly and Ones (2010) pooled 263 samples covering more than 44,000 rated individuals and found that informant ratings predict academic achievement and job performance about as well as self-ratings, and better for some outcomes. Knowing someone well mattered more for accuracy than simply interacting with them often, a pattern that held most clearly for traits that are hard to observe from the outside, such as emotional stability.

That does not mean self-reports are unreliable. It means two honest vantage points on the same person tend to agree, and each adds something the other misses. A validity case for a trait is stronger when it can show both.

How do you check a test's validity in five minutes?

Most people cannot run a replication study before taking a test, but a short check catches the obvious problems. Look for five things: a named trait model with published research behind it, not just a description written for the site itself; facets reported under each broad dimension, a sign the content was built to cover the trait rather than to produce one number; a score placed against other people as a percentile rather than sorted into a small number of boxes; real studies named for any claim about what a score predicts; and a plain statement of what the test is not, such as a tool for identifying medical or psychological conditions.

How accurate are personality tests goes further into what one sitting can and cannot tell you, and are personality tests reliable covers the consistency side of the same coin.

How does Wellington measure validity?

Wellington is a science-based personality assessment from Therabot Labs LLC. It measures personality as continuous traits, built on the Big Five and HEXACO models, and writes the results back as a personal report rather than a type.

The free Snapshot measures six dimensions, the Big Five and Honesty-Humility, with 60 questions. The Wellington Membership extends that to 90 traits, including every Big Five facet, so the item set covers each dimension's range rather than one slice of it, and continues to 255 traits across eight domains.

Every trait is reported as a percentile against the reference sample. The three bands (Quiet below the 30th percentile, Balanced to the 70th, High above) and the Portrait names exist to make the result easier to talk about; the percentile is the measurement, and it is always shown alongside. Our norms are provisional until the norming sample is complete, and we say so rather than presenting an early percentile as a finished one.

Wellington is a wellness tool for reflection, not a medical or psychological diagnosis, and it does not replace care from a qualified professional. It is not designed for hiring, and you can export or delete your data at any time. Read the methodology for how the scores are built.

Questions people ask

What is the difference between reliability and validity?
Reliability is consistency: whether a test gives a similar answer on a retake or across similar items. Validity is accuracy: whether that answer measures what it claims to measure. A classic account of both, and how they relate through measurement error, is Nunnally and Bernstein (1994). A test needs to be reliable before it can be valid, but being reliable alone is not enough.
Can a test be reliable but not valid?
Yes. A bathroom scale that reads five pounds heavy every time is highly reliable, since it gives the same answer over and over, and not valid, since that answer is wrong. Personality items can behave the same way, which is why researchers check content, construct, convergent and predictive validity separately rather than assuming consistency proves accuracy.
Is the Big Five a valid model of personality?
The structural case is strong. The same five-factor pattern recurs across decades of lexical research (Goldberg, 1993) and across self-report and observer data using different instrument types (McCrae & Costa, 1987). Predictive validity adds to the case: the traits forecast outcomes such as occupational attainment about as well as socioeconomic status or cognitive ability (Roberts et al., 2007), and most of that evidence held up under a preregistered replication (Soto, 2019).
Do personality traits actually predict anything in real life?
Yes, modestly and consistently rather than dramatically. Traits predicted mortality, divorce and occupational attainment about as well as socioeconomic status and cognitive ability across prospective longitudinal studies (Roberts et al., 2007). A later replication project found that 87 percent of 78 published trait-outcome links replicated in the expected direction, at roughly 77 percent of the original strength (Soto, 2019).

Sources

Peer-reviewed sources for the claims above. Wellington's own reliability figures will be published once the norming sample is complete.

  1. Nunnally, J. C., & Bernstein, I. H. (1994). Psychometric theory (3rd ed.). McGraw-Hill.
  2. Goldberg, L. R. (1993). The structure of phenotypic personality traits. American Psychologist, 48(1), 26–34.
  3. McCrae, R. R., & Costa, P. T. (1987). Validation of the five-factor model of personality across instruments and observers. Journal of Personality and Social Psychology, 52(1), 81–90.
  4. Roberts, B. W., Kuncel, N. R., Shiner, R., Caspi, A., & Goldberg, L. R. (2007). The power of personality: The comparative validity of personality traits, socioeconomic status, and cognitive ability for predicting important life outcomes. Perspectives on Psychological Science, 2(4), 313–345.
  5. Soto, C. J. (2019). How replicable are links between personality traits and consequential life outcomes? The Life Outcomes of Personality Replication project. Psychological Science, 30(5), 711–727.
  6. Connelly, B. S., & Ones, D. S. (2010). An other perspective on personality: Meta-analytic integration of observers' accuracy and predictive validity. Psychological Bulletin, 136(6), 1092–1122.

Read next

See your own traits, not a type.

The free Snapshot takes about seven minutes and gives you your personality card and five dimensions. No credit card, no type.

One trait a week

Not ready to take the test? Read one trait a week.

A real page from the report, one practice attached, every week. It is the easiest way to see whether Wellington reads people the way you think it should.