Evidence
Is the Big Five scientifically valid?
Yes. The Big Five is the most replicated finding in personality psychology. The same five broad dimensions appear when people describe themselves or others in dozens of languages, scores are stable across decades, roughly forty percent of the variation is heritable, and the traits predict health, relationships and work outcomes about as well as socioeconomic status or measured intelligence.
Last updated September 16, 2026.
Where does the Big Five come from?
The Big Five did not begin as a theory. It began as a pattern in language. When researchers collected the thousands of trait words in the dictionary, asked people to rate themselves and others on them, and analysed which words travelled together, five broad clusters kept appearing: Extraversion, Agreeableness, Conscientiousness, Emotional Stability (its reverse is called Neuroticism) and Openness to Experience (Goldberg, 1990). That structure came out of decades of lexical work rather than any single study (Goldberg, 1993). Questionnaires built from a different direction found the same five, most visibly in the NEO inventories, which measure five domains with six facets each (Costa & McCrae, 1992).
The convergence is the striking part. By 1990 several research programmes that had started from different data, different samples and different theoretical commitments had arrived at the same five-factor solution, and a review that year described the emergence of the five-factor model as the structure the field had converged on (Digman, 1990). Lexical researchers counting adjectives and questionnaire builders writing statements about behaviour were not trying to confirm each other's work. They ended up describing the same space.
That did not settle everything. The five factors carried different names in different traditions, so one group's Culture, another's Intellect and a third's Openness were the same dimension under three labels. John and Srivastava (1999) set the major instruments side by side and gave the model a shared vocabulary. Modern inventories inherit it, which is why a score from one can be read against a score from another.
What is the evidence that it is real?
- It replicates across cultures. Observer ratings collected in fifty cultures reproduced the five-factor structure (McCrae & Terracciano, 2005).
- It is stable. A review of longitudinal studies found that people keep their rank order on the traits to a high degree across adulthood, rising with age (Roberts & DelVecchio, 2000). Traits change slowly, and the changes are the interesting part.
- It is partly heritable. A meta-analysis of twin and family studies put the heritability of personality at about forty percent (Vukasović & Bratko, 2015).
- It predicts outcomes. Across large longitudinal studies, personality traits predicted mortality, divorce and occupational attainment about as well as socioeconomic status and cognitive ability (Roberts et al., 2007).
- It has a hierarchy. Under each broad dimension sit narrower facets that add prediction beyond the domain score, which is why modern inventories measure both (Soto & John, 2017).
| The claim | What was actually tested | Where it comes from |
|---|---|---|
| Five dimensions keep appearing | Trait adjectives factor analysed in large samples; separate questionnaire programmes converging | Goldberg (1990); Digman (1990) |
| It is not an artefact of English | Self-descriptions from 56 nations in 28 languages; observer ratings from 50 cultures | Schmitt et al. (2007); McCrae & Terracciano (2005) |
| Scores hold their order over years | The same people retested, from childhood to old age | Roberts & DelVecchio (2000); Specht et al. (2011) |
| Part of the variation is inherited | Twin and family designs, pooled across studies | Vukasović & Bratko (2015) |
| The traits predict real outcomes | Cohort data on mortality, divorce and attainment; a preregistered replication | Roberts et al. (2007); Soto (2019) |
None of this shows that five is a magic number, that the dimensions are causes rather than descriptions, or that a score settles anything about one person on one afternoon. It shows that the map is drawn from the territory and keeps redrawing the same way.
Does the Big Five hold up outside English-speaking countries?
Mostly, with honest exceptions. The largest single test of this administered a Big Five inventory translated into 28 languages to samples from 56 nations, and the five-factor structure was recovered in the great majority of them (Schmitt et al., 2007). A separate project asked people in fifty cultures to rate someone they knew well rather than themselves, which removes self-presentation from the picture, and the same five factors appeared again (McCrae & Terracciano, 2005).
The exceptions are worth stating. The authors flag limitations in the data set themselves, including how the samples were recruited, and translated inventories and unfamiliar response formats push in the same direction (Schmitt et al., 2007). The careful conclusion is that the five dimensions are widely, not universally, recoverable, and that mean scores from different countries should not be compared casually.
A structural caveat travels the other way. Lexical studies run from scratch in several languages, rather than by translating an English inventory, tend to produce six factors, with sincerity, fairness and modesty separating out (Ashton & Lee, 2007). That is an argument for measuring more than five, not fewer.
What does the Big Five predict?
This is the question that decides whether a model is useful rather than merely tidy. A review covering individual, interpersonal and institutional outcomes found that Big Five traits relate to happiness and health, to the quality of relationships with peers, family and partners, and to occupational choice, job performance and community involvement (Ozer & Benet-Martínez, 2006). The effects look small in isolation and accumulate over a life.
The obvious worry is that such findings come from decades of analyses that were not planned in advance. That worry has been tested directly. A preregistered project re-examined a large set of published trait-outcome associations in a new sample, and most replicated in the same direction, at effect sizes somewhat smaller than first reported (Soto, 2019). Some did not, and the project named them.
Read all of it as averages across many people rather than as forecasts about one. An effect that shifts the odds a little, every week, for forty years, is not a small effect; it is a slow one. Trait by trait the picture is more specific: Conscientiousness carries the work and health findings, Emotional Stability most of the well-being findings, Extraversion the leadership findings. The glossary sets out each dimension, and the facets page goes a level below them.
Is the Big Five stable over a lifetime?
Two different questions hide inside that one, and the answers differ. Rank-order stability asks whether the most conscientious person in a group stays near the top of it years later. A meta-analysis of longitudinal studies found that it largely does, and that this consistency rises steadily with age, from moderate in childhood to high in the decades after fifty (Roberts & DelVecchio, 2000).
Mean-level change asks whether people as a group shift. They do, in a direction researchers call the maturity principle: Conscientiousness, Emotional Stability and social dominance rise through adulthood, mostly between twenty and forty, Agreeableness changes mainly in old age, and social vitality rises in adolescence before declining late (Roberts et al., 2006). A large German study following about fifteen thousand adults added that rank-order stability itself peaks in midlife rather than climbing forever, and that particular life events move some traits and leave others untouched (Specht et al., 2011).
Put together: you are recognisably the same person at fifty that you were at twenty, and you are measurably not the same person. Both halves come from the same data.
What are the limits?
The Big Five is a description of how trait words and self-reports organise, not an explanation of why. It says little about motives, narratives or the situations that bring a trait out. Its measures are self-reports, which are honest for most people most of the time but not immune to mood or self-presentation. And five may not be the whole story at the broad level: lexical studies in several languages find a sixth factor, Honesty-Humility, which is why the HEXACO model exists and why Wellington measures six dimensions rather than five (Ashton & Lee, 2007).
None of these limits makes the model invalid. They describe what kind of thing it is: a reliable map of broad differences between people, best read alongside a life rather than instead of one.
Two further cautions belong here. A domain score hides the facets underneath it, so two people at the same percentile on Conscientiousness can differ on whether that is orderliness or industriousness (Soto & John, 2017). And a percentile is a comparison rather than a quantity, so it depends on the sample you are placed against. How to read personality test scores covers what a band means, and how accurate are personality tests covers how much weight one sitting can bear.
How does Wellington use it?
Wellington is a science-based personality assessment from Therabot Labs LLC. It measures personality as continuous traits, built on the Big Five and HEXACO models, and writes the results back as a personal report rather than a type.
The free Snapshot measures the Big Five and Honesty-Humility with 60 questions. The Wellington Membership adds every facet beneath them and the 24 VIA character strengths, then continues to 255 traits across eight domains, adding emotional functioning, relationships, stress and resilience, meaning and values, lifestyle and self-concept. Read the methodology for how the scores are made. Wellington is a wellness tool for reflection, not a medical or psychological diagnosis, and it does not replace care from a qualified professional. It is not designed for hiring, and you can export or delete your data at any time.
The Snapshot takes about 7 minutes. Each dimension is the average of ten statements rated on a five-point accuracy scale, placed against a reference sample as a percentile, and reported in one of three bands: Quiet below the 30th percentile, Balanced to the 70th, High above it. Those norms are provisional until the norming sample is complete, and the report says so. The bands, and the Portrait name that summarises them with the membership, exist to make the result easier to talk about; the percentile is the measurement, and it is always shown alongside. The first ten questions are answered before anything is asked of you; after question ten you sign in with an email so your answers are saved and your results can be shown. No card is needed.
In the written report, every trait carries three pages, one per band, prepared in advance by the writing and psychology team, and your answers choose which one you read. The 90 traits with the membership are chosen so the evidence above applies to them: broad dimensions with a long literature, the facets beneath them, and the character strengths measured alongside.
Questions people ask
- Is the Big Five better than the MBTI?
- As measurement, yes. The MBTI's four scales relate to four of the Big Five dimensions, so they cover similar ground (McCrae & Costa, 1989). The difference is what happens after the scoring. The MBTI rounds each scale to one side of a midpoint, has no scale for the dimension the Big Five calls Emotional Stability, and produces a four-letter formula that supports fewer inferences than its users draw from it, which is why reviews of the instrument counsel caution (Pittenger, 2005). A percentile keeps the degree a letter throws away.
- Can personality change?
- Yes, slowly. People largely keep their rank order relative to others, and that consistency rises with age (Roberts & DelVecchio, 2000). Average levels still move: conscientiousness and emotional stability rise across adulthood, most steeply between twenty and forty, and agreeableness changes mainly in old age (Roberts et al., 2006). Deliberate change is possible and modest, and it comes from repeated behaviour rather than from insight. Wellington's weekly modules and its retests in year two are built for that timescale.
- Why does Wellington measure six dimensions?
- Because lexical studies run from scratch in several languages keep producing six broad factors rather than five, with sincerity, fairness, greed avoidance and modesty separating out as Honesty-Humility (Ashton & Lee, 2007). The five-factor solution folds that content into Agreeableness, where it is easy to miss. Measuring it on its own adds information about how a person behaves when there is an advantage to be taken.
- Did the Big Five survive the replication crisis?
- Better than most areas of psychology. A preregistered project retested a large set of published trait-outcome associations in a fresh sample, and most replicated in the same direction with somewhat smaller effects (Soto, 2019). The structural findings rest on very large samples across many languages.
- Is the Big Five the same in every country?
- Broadly, not perfectly. Self-descriptions from 56 nations and observer ratings from 50 cultures recovered the same five dimensions in most samples (Schmitt et al., 2007; McCrae & Terracciano, 2005). Fit was weaker in some regions, and translated inventories may be part of the reason.
Sources
Peer-reviewed sources for the claims above. Wellington's own reliability figures will be published once the norming sample is complete.
- Goldberg, L. R. (1990). An alternative "description of personality": The Big-Five factor structure. Journal of Personality and Social Psychology, 59(6), 1216–1229.
- Goldberg, L. R. (1993). The structure of phenotypic personality traits. American Psychologist, 48(1), 26–34.
- Digman, J. M. (1990). Personality structure: Emergence of the five-factor model. Annual Review of Psychology, 41, 417–440.
- John, O. P., & Srivastava, S. (1999). The Big Five trait taxonomy: History, measurement, and theoretical perspectives. In L. A. Pervin & O. P. John (Eds.), Handbook of personality: Theory and research (2nd ed., pp. 102–138). Guilford Press.
- Costa, P. T., & McCrae, R. R. (1992). Revised NEO Personality Inventory (NEO PI-R) and NEO Five-Factor Inventory (NEO-FFI) professional manual. Psychological Assessment Resources.
- McCrae, R. R., & Terracciano, A. (2005). Universal features of personality traits from the observer's perspective: Data from 50 cultures. Journal of Personality and Social Psychology, 88(3), 547–561.
- Schmitt, D. P., Allik, J., McCrae, R. R., & Benet-Martínez, V. (2007). The geographic distribution of Big Five personality traits: Patterns and profiles of human self-description across 56 nations. Journal of Cross-Cultural Psychology, 38(2), 173–212.
- Roberts, B. W., & DelVecchio, W. F. (2000). The rank-order consistency of personality traits from childhood to old age: A quantitative review of longitudinal studies. Psychological Bulletin, 126(1), 3–25.
- Roberts, B. W., Walton, K. E., & Viechtbauer, W. (2006). Patterns of mean-level change in personality traits across the life course: A meta-analysis of longitudinal studies. Psychological Bulletin, 132(1), 1–25.
- Specht, J., Egloff, B., & Schmukle, S. C. (2011). Stability and change of personality across the life course: The impact of age and major life events on mean-level and rank-order stability of the Big Five. Journal of Personality and Social Psychology, 101(4), 862–882.
- Vukasović, T., & Bratko, D. (2015). Heritability of personality: A meta-analysis of behavior genetic studies. Psychological Bulletin, 141(4), 769–785.
- Roberts, B. W., Kuncel, N. R., Shiner, R., Caspi, A., & Goldberg, L. R. (2007). The power of personality: The comparative validity of personality traits, socioeconomic status, and cognitive ability for predicting important life outcomes. Perspectives on Psychological Science, 2(4), 313–345.
- Ozer, D. J., & Benet-Martínez, V. (2006). Personality and the prediction of consequential outcomes. Annual Review of Psychology, 57, 401–421.
- Soto, C. J. (2019). How replicable are links between personality traits and consequential life outcomes? The Life Outcomes of Personality Replication project. Psychological Science, 30(5), 711–727.
- Soto, C. J., & John, O. P. (2017). The next Big Five Inventory (BFI-2): Developing and assessing a hierarchical model with 15 facets to enhance bandwidth, fidelity, and predictive power. Journal of Personality and Social Psychology, 113(1), 117–143.
- Ashton, M. C., & Lee, K. (2007). Empirical, theoretical, and practical advantages of the HEXACO model of personality structure. Personality and Social Psychology Review, 11(2), 150–166.
- McCrae, R. R., & Costa, P. T. (1989). Reinterpreting the Myers-Briggs Type Indicator from the perspective of the five-factor model of personality. Journal of Personality, 57(1), 17–40.
- Pittenger, D. J. (2005). Cautionary comments regarding the Myers-Briggs Type Indicator. Consulting Psychology Journal: Practice and Research, 57(3), 210–221.
Read next
See your own traits, not a type.
The free Snapshot takes about seven minutes and gives you your personality card and five dimensions. No credit card, no type.