Skip to content
Seven minutes. A new way to see yourself.Get your free results →
Wellington

Guide

Psychology tests for personality: which kinds to trust

Psychology uses several quite different kinds of personality test: self-report trait inventories, observer ratings, projective methods, type indicators, strengths surveys and ability tests. They differ enormously in how much evidence stands behind them. The self-report trait inventories built on replicated models are the best supported for describing ordinary personality.

Last updated September 16, 2026.

What is a psychology test for personality?

A psychological test is a standardised procedure for putting a number on something about a person. Standardised is the load-bearing word. Everybody gets the same items in the same order with the same response options, the answers are scored by a fixed rule, and the result is interpreted against a comparison group whose composition is described somewhere you can read it.

That definition rules out a good deal of what gets called a psychology test online. A set of questions written in an afternoon, scored by an undocumented rule and compared with nothing, is a quiz. It may still be enjoyable. It is not measurement, and the difference matters as soon as you want to act on the result.

Two uses are also worth separating. Describing where an ordinary person sits on ordinary traits is one job. Assessing whether someone needs care is a different job, done by qualified professionals with different instruments and a conversation. Everything below is about the first.

What families of personality test exist?

Six families cover almost everything you will meet. They ask different questions, and their evidence bases are not comparable.

Self-report trait inventories ask you to rate statements about yourself and report continuous scores on dimensions that emerged from decades of lexical and questionnaire research (Goldberg, 1993). The Big Five and HEXACO inventories are the main examples, and they are the best supported instruments for describing normal-range personality.

Observer reports ask people who know you to rate the same statements about you. This is not a fallback for when self-report is unavailable. A meta-analytic review of 263 samples found that observer ratings predict behaviour about as well as self-ratings, and better for some outcomes, including academic achievement and job performance, and that knowing someone well matters more for accuracy than seeing them often (Connelly & Ones, 2010). Self and other simply know different things.

Projective methods present an ambiguous stimulus, an inkblot or an unfinished sentence, on the theory that what you supply will reveal something you would not report directly. They have a long history and are still in use in some settings. For measuring normal-range personality traits, the supporting evidence is much weaker than for well-built self-report inventories, and results depend heavily on the scoring system and on the person applying it.

Type indicators sort people into categories: sixteen types, nine types, four colours. The best-studied example is the MBTI, whose scales do relate to real trait dimensions but whose four-letter type formula does not support the inferences commonly drawn from it (Pittenger, 2005). The vocabulary is memorable; the categories are not stable enough to carry decisions.

Strengths surveys ask what you are like at your best rather than how you differ from the average. The most developed is the classification of 24 character strengths under six virtues assembled by Peterson and Seligman (Peterson & Seligman, 2004), which has been surveyed across many countries (McGrath, 2015).

Ability tests score performance rather than self-description. Items have better and worse answers, scored against expert or consensus judgement. Ability-based emotional intelligence tests are the main example in the personality space, and their designers argue that an intelligence has to be demonstrated rather than claimed (Mayer et al., 2008).

How do the families compare?

The main families of personality measure and where their evidence stands
FamilyWhat it measuresHow it worksHow strong the evidence is
Self-report trait inventoryPosition on continuous dimensions such as the Big FiveYou rate statements about yourselfStrongest for ordinary personality; the structure replicates across languages and samples
Observer reportThe same dimensions, as seen by people who know youSomeone else rates the statements about youStrong; ratings predict behaviour about as well as self-ratings, and better for some outcomes (Connelly & Ones, 2010)
Strengths surveyWhich valued qualities you draw on mostYou rate statements, and strengths are rankedReasonable; a published classification surveyed widely (Peterson & Seligman, 2004)
Ability testPerformance on tasks, for example reasoning about emotionItems with better and worse answersMixed and still debated; depends on how the correct answer is defined
Type indicatorMembership of a categoryAnswers are summed and cut into typesWeak for the types themselves; the underlying scales track real traits (Pittenger, 2005)
Projective methodThemes in what you produce from an ambiguous promptA trained scorer codes your responsesWeak for normal-range trait measurement; heavily dependent on the scoring system

What makes a psychological test trustworthy?

Three properties, and they are checkable before you take anything.

  • Reliability. The test gives the same answer when nothing about you has changed, and its items agree with each other. Reliability rises with the number of items per scale, which is the oldest result in the field (Nunnally & Bernstein, 1994). Very short scales pay for their brevity: an evaluation of short Big Five measures found they substantially understate the part traits play in important behaviours (Credé et al., 2012).
  • Validity. The scores relate to things outside the test: to behaviour, to how other people describe you, to outcomes recorded later. A test can be perfectly reliable and measure nothing useful.
  • Norms. Your raw score means nothing on its own. It has to be set against a described comparison group, and a trustworthy test says who that group was and how it was gathered.

Two further things are worth checking. A good test reports parts rather than one blended number, so you can see where a score comes from. And a good result says something you could have disagreed with, rather than something that would fit anybody who read it. A result that fits everyone has not measured anyone.

What can a psychology test not tell you?

It cannot tell you whether you are healthy or unwell. A personality questionnaire describes ordinary differences between ordinary people, on traits where every position on the scale is a normal place to be. Questions about your health belong with a qualified professional who can talk to you, not with a questionnaire and a score.

It cannot tell you who to hire. Answers that are honest and useful for reflection behave differently when the person answering has a reason to look good, and selection is a regulated field with its own standards. Nor can it tell you what to do with your life. Traits describe tendencies, not destinations, and the link between a trait and any single decision is loose.

And it cannot tell you more than you put in. Self-report is limited by self-knowledge, which is real but partial. That is an argument for asking someone who knows you well to answer the same questions about you, and for treating a divergence between the two as information rather than as an error (Connelly & Ones, 2010).

How does Wellington fit?

Wellington is a science-based personality assessment from Therabot Labs LLC. It measures personality as continuous traits, built on the Big Five and HEXACO models, and writes the results back as a personal report rather than a type.

Wellington sits in the first family, and it uses the second and third where they help. The free Snapshot is 60 questions, about 7 minutes, and measures six broad dimensions: Extraversion, Agreeableness, Conscientiousness, Emotional Stability, Openness to Experience, Honesty-Humility. Each score is the average of several statements rated on a five-point accuracy scale, set against a reference sample as a percentile and reported in one of three bands, Quiet below the 30th percentile, Balanced to the 70th, High above. Our norms are provisional until the norming sample is complete. The bands and the Portrait names exist to make the result easier to talk about; the percentile is the measurement, and it is always shown alongside. The first ten questions are answered before anything is asked; after question ten you sign in with an email so your answers are saved and your results can be shown. There is no card.

The Wellington Membership, $4.99 a month, adds the Portrait, 336 questions in 4 chapters, to reach 90 traits, including every Big Five facet and all 24 VIA character strengths (Peterson & Seligman, 2004), with domain overviews, a strengths constellation, a growth plan and a downloadable PDF, then adds a short weekly module over 104 weeks toward all 255 traits across eight domains, with retests in year two. You can cancel any time.

In the written report, every trait has three pages, one per band, written in advance by the writing and psychology team, and your answers choose which page you read. Wellington is made by a psychologist and a psychiatrist at Therabot Labs LLC. Wellington is a wellness tool for reflection, not a medical or psychological diagnosis, and it does not replace care from a qualified professional. It is not for hiring, and you can export or delete your data at any time. The methods page sets out how the items were written and scored.

Questions people ask

What is the most scientifically valid personality test?
An inventory built on the Big Five or HEXACO models, reporting continuous scores with published reliability and a described comparison sample. The BFI-2, the IPIP-NEO and the HEXACO-PI-R are free and well documented. What makes them strong is the model underneath and the number of questions per trait, not the brand on the page.
Are online psychology tests real tests?
Some are, some are not. A real one names its model, asks several questions per scale, scores by a fixed rule and compares your answers with a described sample. A quiz that returns a label with no explanation of how it got there is entertainment. Both can be online; the difference is in the documentation, not the medium.
What is the difference between a personality test and an ability test?
A personality test asks what you are typically like, and there are no right answers. An ability test asks what you can do, and the items are scored against a standard. Ability-based emotional intelligence tests are the crossover case, treating emotional understanding as a skill to be demonstrated rather than a self-description (Mayer et al., 2008).
Should someone else take the test about me?
It is a genuinely good idea. Observer ratings predict behaviour about as well as self-ratings, and better for some outcomes such as job and academic performance, with accuracy depending more on how well the rater knows you than on how often they see you (Connelly & Ones, 2010). You know your inner life best; people who know you often read your outward habits more accurately. Disagreements between the two are informative.

Sources

Peer-reviewed sources for the claims above. Wellington's own reliability figures will be published once the norming sample is complete.

  1. Goldberg, L. R. (1993). The structure of phenotypic personality traits. American Psychologist, 48(1), 26–34.
  2. Connelly, B. S., & Ones, D. S. (2010). An other perspective on personality: Meta-analytic integration of observers' accuracy and predictive validity. Psychological Bulletin, 136(6), 1092–1122.
  3. Pittenger, D. J. (2005). Cautionary comments regarding the Myers-Briggs Type Indicator. Consulting Psychology Journal: Practice and Research, 57(3), 210–221.
  4. Peterson, C., & Seligman, M. E. P. (2004). Character Strengths and Virtues: A Handbook and Classification. Oxford University Press and the American Psychological Association.
  5. McGrath, R. E. (2015). Character strengths in 75 nations: An update. The Journal of Positive Psychology, 10(1), 41–52.
  6. Mayer, J. D., Salovey, P., & Caruso, D. R. (2008). Emotional intelligence: New ability or eclectic traits? American Psychologist, 63(6), 503–517.
  7. Nunnally, J. C., & Bernstein, I. H. (1994). Psychometric theory (3rd ed.). McGraw-Hill.
  8. Credé, M., Harms, P., Niehorster, S., & Gaye-Valentine, A. (2012). An evaluation of the consequences of using short measures of the Big Five personality traits. Journal of Personality and Social Psychology, 102(4), 874–888.

Read next

See your own traits, not a type.

The free Snapshot takes about seven minutes and gives you your personality card and five dimensions. No credit card, no type.

One trait a week

Not ready to take the test? Read one trait a week.

A real page from the report, one practice attached, every week. It is the easiest way to see whether Wellington reads people the way you think it should.