Skip to content
Seven minutes. A new way to see yourself.Get your free results →
Wellington

Evidence

A short history of personality tests, from word lists to the Big Five

Personality testing grew out of a single idea: that the differences between people worth talking about could be named, counted and measured rather than left to impression. Allport gave the trait its modern definition in 1937, lexical researchers spent the following decades sorting trait words into five recurring factors, questionnaire studies found the same five from a different direction, and later work refined, shortened and made those tools public.

Last updated September 16, 2026.

What is the trait concept, and where does this history start?

Most histories of personality testing start with the same question: is a personality made of stable traits at all, or only of behaviour that shifts with the situation? Allport's (1937) monograph is the conventional starting point because it argued for the trait as psychology's basic unit for describing a person, a consistent way of thinking, feeling or acting that shows up across situations rather than being invented anew each time. Before a trait could be measured, someone had to argue it was a real, nameable thing, and Allport's book is the standard citation for that argument.

That starting point matters because everything that follows depends on it. A questionnaire, a word list or a factor analysis only makes sense as a way of measuring something if the something is assumed to be stable enough to measure. The rest of this history is largely the story of researchers testing whether that assumption held up, and how many traits it takes to describe a person once it does.

How did the lexical approach turn word lists into five factors?

One answer to how many traits exist came from language rather than theory. The lexical approach assumes that important differences between people accumulate words over generations of people describing each other, so a language's trait vocabulary is a rough map of what its speakers have found worth noticing. The lexical hypothesis covers that assumption and its limits in full; the short version here is what it produced.

Goldberg's (1990) three-study programme is the clearest demonstration. Large sets of English trait adjectives were rated by self-report and by peers, then run through factor analysis, a technique that groups items which rise and fall together into a smaller number of underlying dimensions. Across different samples and different factor-analytic procedures, the same five-factor structure kept reappearing, and no sixth factor generalised across the studies. That repetition, rather than any single result, is what gave the five factors their standing.

How did the field come to treat five factors as the consensus?

A structure recovered by one research programme is not yet a field-wide consensus. Goldberg (1993) supplies the history of how it became one, tracing the idea through Galton, Thurstone, Cattell, and Tupes and Christal, and arguing that the model's acceptance owes a great deal to its critics: each attempt to replace it with something else failed to dislodge it. On that account, the Big Five is less a single discovery than the survivor of decades of people trying to prove it wrong, described from within the field rather than reported as a study with its own new data.

Did the questionnaire tradition find the same five factors?

Word lists are one route to a trait structure. Questionnaires, where people rate themselves against written statements rather than single adjectives, are another, and it mattered whether the two routes agreed. McCrae and Costa (1987) tested this directly, comparing adjective factors against questionnaire scales using both self-reports and peer ratings of hundreds of adults. The same five factors showed up in both formats, and self-reports agreed substantially with peer ratings on all five, evidence that the structure was not an artefact of asking people to rate single words.

That convergence fed directly into instrument-building. Costa and McCrae's (1992) NEO Personality Inventory and its short form, the NEO-FFI, documented five domains, each built from six facets, as a professional instrument other researchers and clinicians could use rather than a one-off research finding. The NEO family became one of the standard questionnaire measures of the five-factor model for the following decades.

What did later reviews conclude about the popular type tests?

While the five-factor research was converging, a separate tradition of type-based tests, most visibly the Myers-Briggs Type Indicator, was becoming popular in workplaces and counselling. Pittenger's (2005) review of the psychometric evidence found that the four-letter type formula does not support the inferences commonly drawn from it, and urged real caution about the claims made for it in professional settings. The MBTI's popularity and its evidentiary standing turned out to be two different questions.

Part of why a vague type description can feel so accurate is a separate finding with its own history. Forer (1949) gave students an identical, generic personality sketch, told each of them it had been prepared just for them, and found they rated it as strikingly accurate on average, a classroom demonstration that people readily accept broad, flattering descriptions as personally true. That pattern, now usually called the Forer or Barnum effect, is a caution worth carrying into how any personality description, including a scientific one, gets read.

How did personality measurement become public domain, and what came after?

For much of this history, the best trait measures were commercial instruments controlled by their publishers. Goldberg and six other measurement researchers argued for a different model with the International Personality Item Pool, a public-domain bank of items intended as a real alternative to proprietary inventories rather than a rough substitute for them (Goldberg et al., 2006). Items built in that tradition could be reused, adapted and studied without a licence, which changed who could build a trait measure at all.

The questionnaire tradition kept developing alongside it. Soto and John's (2017) BFI-2 organised the five domains into a hierarchy of 15 facets, with 60 items in all, 12 per domain and 4 per facet, validated against both self-reports and peer reports, adding the bandwidth of facet-level detail without abandoning the five-factor structure the earlier decades had converged on. The authors make the instrument available at no cost for research and educational use.

Is there a sixth factor beyond the Big Five?

The five-factor structure was never the final word, even among the researchers who established it. Ashton and Lee (2007) argued for a six-factor alternative, HEXACO, built on a lexically derived structure that adds Honesty-Humility to versions of Extraversion, Agreeableness, Conscientiousness, Emotionality and Openness. They point to phenomena a five-factor model does not explain well, including patterns around altruism and sex differences in traits, as reasons the sixth factor earns its place rather than merely restating the other five. What is the five-factor model covers how the five- and six-factor accounts relate.

What are the milestones in this history, and what should it teach you about choosing a test today?

Ten milestones in the history of personality testing, in order
YearMilestoneWhat it contributed
1937Allport's trait monographThe standard case for the trait as a stable, measurable unit of personality
1949Forer's classroom demonstrationThe finding that people accept a generic description as personally accurate, a caution for reading any personality report
1987McCrae and Costa's validation studyThe same five factors confirmed across adjective and questionnaire formats, and across self and peer ratings
1990Goldberg's lexical studiesThe five-factor structure recovered repeatedly from English trait adjectives
1992Costa and McCrae's NEO PI-R manualA widely used questionnaire measure of the five domains, each with six facets
1993Goldberg's history of the taxonomyAn account of how the five-factor model became the field's working consensus
2005Pittenger's review of the MBTIEvidence that a popular four-letter type formula does not support its common uses
2006The International Personality Item PoolA public-domain alternative to commercial trait inventories
2007Ashton and Lee's HEXACO reviewA six-factor alternative adding Honesty-Humility to the five-factor structure
2017Soto and John's BFI-2A validated 15-facet hierarchy beneath the five domains, free for research and education

Read in order, this history points to a few practical lessons for choosing a test today. The trait concept behind every instrument here has been argued for, not assumed, since 1937. The five-factor structure did not come from a single clever study; it came from multiple methods, adjectives and questionnaires, self-reports and peer reports, converging on the same answer, which is a stronger kind of evidence than any one paper could give. Popularity and evidentiary standing are separate questions, as the MBTI's history shows, and a description that feels accurate is not proof that it is, given how readily people accept a generic sketch as personal (Forer, 1949). And the field kept revising itself long after the Big Five settled in, adding facets, shortening item sets, making items public, and arguing for a sixth factor, which is what an active scientific tradition looks like rather than a finished one.

Best personality tests applies these same distinctions to instruments available today, and what is personality psychology covers the wider field this history sits inside.

How does Wellington fit into this history?

Wellington is a science-based personality assessment from Therabot Labs LLC. It measures personality as continuous traits, built on the Big Five and HEXACO models, and writes the results back as a personal report rather than a type.

Wellington sits in the questionnaire tradition this history describes, not the type tradition. The free Snapshot asks 60 questions, takes about 7 minutes, and reports 6 dimensions, the five factors this history traces plus Honesty-Humility from the HEXACO work: Extraversion, Agreeableness, Conscientiousness, Emotional Stability, Openness to Experience, Honesty-Humility. Each question is a statement about behaviour rated on a five-point accuracy scale, in the spirit of the questionnaire studies above rather than a single adjective checklist. The statements are real public-domain IPIP items and public-domain items written in the same style, curated for a wellness setting, so the Snapshot sits in the public-domain tradition this history describes. The first ten questions are answered before anything else is asked; after question ten you sign in with an email so your answers are saved and your results can be shown.

Every trait is reported as a percentile against the reference sample. The three bands (Quiet below the 30th percentile, Balanced to the 70th, High above) and the Portrait names built from them exist to make the result easier to talk about; the percentile is the measurement, and it is always shown alongside. Norms are provisional until the norming sample is complete. The Wellington Membership, $4.99 a month, extends this to 90 traits, including every Big Five facet and all 24 VIA character strengths, and carries the method toward 255 traits. Wellington is a wellness tool for reflection, not a medical or psychological diagnosis, and it does not replace care from a qualified professional. It is not built for hiring, and you can export or delete your data at any time.

Questions people ask

Who invented personality tests?
No single person. Allport's (1937) monograph gave the trait its standard modern definition, and the tests that followed were built by many researchers over decades: Goldberg's (1990, 1993) lexical studies, McCrae and Costa's (1987) questionnaire work, and Costa and McCrae's (1992) NEO inventories among them. Personality testing is the product of a long, argumentative research tradition, not one inventor.
When was the Big Five created?
There is no single founding date. Goldberg's (1990) three-study lexical programme is the clearest single demonstration of the five-factor structure, and McCrae and Costa (1987) confirmed the same structure from a questionnaire angle around the same time, but Goldberg's (1993) own history traces the idea back through Galton, Thurstone, Cattell, and Tupes and Christal, decades earlier.
Why did the MBTI become so popular if the evidence is limited?
Popularity and evidentiary standing are separate questions. Pittenger's (2005) review found the four-letter type formula does not support the inferences often drawn from it, but a type label is simple, memorable and flattering to discuss, and generic descriptions tend to feel personally accurate even when they are not (Forer, 1949).
Are personality tests today based on the same Big Five research?
Many are. The NEO inventories (Costa & McCrae, 1992) and the BFI-2 (Soto & John, 2017) both measure the five factors this history traces, with facets added beneath them, and the public-domain International Personality Item Pool (Goldberg et al., 2006) let other tests build on the same tradition without a commercial licence.
Is the Big Five the final answer in this history?
No. Ashton and Lee (2007) argue for a six-factor HEXACO model that adds Honesty-Humility to versions of the five factors, built on lexical research that has replicated across several languages. The history is an ongoing argument, not a closed case.

Sources

Peer-reviewed sources for the claims above. Wellington's own reliability figures will be published once the norming sample is complete.

  1. Allport, G. W. (1937). Personality: A psychological interpretation. Holt.
  2. Goldberg, L. R. (1990). An alternative "description of personality": The Big-Five factor structure. Journal of Personality and Social Psychology, 59(6), 1216–1229.
  3. Goldberg, L. R. (1993). The structure of phenotypic personality traits. American Psychologist, 48(1), 26–34.
  4. McCrae, R. R., & Costa, P. T. (1987). Validation of the five-factor model of personality across instruments and observers. Journal of Personality and Social Psychology, 52(1), 81–90.
  5. Costa, P. T., & McCrae, R. R. (1992). Revised NEO Personality Inventory (NEO PI-R) and NEO Five-Factor Inventory (NEO-FFI) professional manual. Psychological Assessment Resources.
  6. Pittenger, D. J. (2005). Cautionary comments regarding the Myers-Briggs Type Indicator. Consulting Psychology Journal: Practice and Research, 57(3), 210–221.
  7. Forer, B. R. (1949). The fallacy of personal validation: A classroom demonstration of gullibility. Journal of Abnormal and Social Psychology, 44(1), 118–123.
  8. Goldberg, L. R., Johnson, J. A., Eber, H. W., Hogan, R., Ashton, M. C., Cloninger, C. R., & Gough, H. G. (2006). The international personality item pool and the future of public-domain personality measures. Journal of Research in Personality, 40(1), 84–96.
  9. Soto, C. J., & John, O. P. (2017). The next Big Five Inventory (BFI-2): Developing and assessing a hierarchical model with 15 facets to enhance bandwidth, fidelity, and predictive power. Journal of Personality and Social Psychology, 113(1), 117–143.
  10. Ashton, M. C., & Lee, K. (2007). Empirical, theoretical, and practical advantages of the HEXACO model of personality structure. Personality and Social Psychology Review, 11(2), 150–166.

Read next

See your own traits, not a type.

The free Snapshot takes about seven minutes and gives you your personality card and five dimensions. No credit card, no type.

One trait a week

Not ready to take the test? Read one trait a week.

A real page from the report, one practice attached, every week. It is the easiest way to see whether Wellington reads people the way you think it should.