Guide
How accurate are online personality tests?
Accuracy has two halves: whether a test gives the same answer twice, and whether its scores relate to anything real. A Big Five or HEXACO inventory with several statements per trait does reasonably well on both. Tests that sort you into a type do worse on the first, and formats read from handwriting have not held up on the second.
Last updated September 16, 2026.
What does an accurate personality test mean?
Accuracy is not one property. Psychometrics splits it into two questions, and a test can pass one and fail the other (Nunnally & Bernstein, 1994).
Reliability is consistency. Does the test give you roughly the same score when you take it again a few weeks later, and do the statements that are meant to measure the same trait actually move together? An unreliable test cannot be valid, because a score that changes for no reason cannot be tracking anything.
Validity is whether the score means what the label says. Does a Conscientiousness score relate to finishing what you start, keeping appointments and the behaviour other people notice in you? Does it predict anything measurable, months later, in someone who was not in the sample used to build the test? Broad traits do predict consequential outcomes, including occupational attainment, relationship stability and mortality, at a level comparable with socioeconomic status and cognitive ability (Roberts et al., 2007), and many specific links hold up when tested again in fresh samples (Soto, 2019).
The honest answer for a good trait inventory is that it is accurate about groups and approximate about individuals. It can place you relative to other people, with a margin of error. It cannot tell you what you will do on Tuesday.
What makes one test more accurate than another?
Three things, in roughly this order of importance.
- A model that has been replicated. The five-factor structure came out of decades of lexical research rather than a single study, and it survived repeated attempts to replace it (Goldberg, 1993). HEXACO adds a sixth factor from a lexical structure that has replicated across cultures (Ashton & Lee, 2007). A model invented for one product has none of that behind it.
- Enough statements per trait. Reliability rises with item count. Four-item and two-item measures exist and are useful when survey space is scarce, and their own authors are clear about the cost in precision (Gosling et al., 2003). Very abbreviated measures also lead researchers to understate how much traits matter, while slightly longer scales are substantially more valid for little extra time (Credé et al., 2012).
- Norms that are described. A score only becomes a percentile once it is placed against a reference sample. A test that will not say who that sample is, or that norms against everyone its advertising happened to reach, is giving you a number with an unstated denominator.
Two things that do not make a test accurate: how well the result describes you on first reading, and how long it took. A description can feel exact and measure nothing.
How accurate is self-report, and what does it miss?
Almost every personality test asks you about yourself, which works better than people expect and has a predictable blind spot. You know more than anyone else about your inner life: how anxious you feel, what you are drawn to, what you are thinking in a meeting. Other people know more about the traits that are visible from outside and carry social stakes, which is why, against behavioural criteria, the self was the best judge of neuroticism-related traits, friends were the best judges of intellect-related traits, and both did equally well on extraversion (Vazire, 2010).
Ratings by people who know you are not merely a second opinion. Observer ratings predict behaviour about as well as self-ratings, and better for some outcomes, including academic achievement and job performance; knowing someone well matters more for accuracy than seeing them often (Connelly & Ones, 2010).
The practical reading is modest. A self-report score is worth having. If it matters, ask two people who know you in different settings what they would have answered, and pay attention to where they disagree with you.
Why do type tests give a different answer on a retake?
Not because their questions are bad. Type inventories in the Myers-Briggs tradition measure scales that map onto four of the Big Five dimensions reasonably well (McCrae & Costa, 1989). The trouble comes after the scoring, when each continuous scale is cut at a midpoint and you are placed on one side of it.
Most people sit near the middle of most scales, so the cut lands where the scores are densest. An ordinary wobble in one answer moves you across the line and changes a letter, and a changed letter changes the whole description. The scores themselves show no sign of the qualitatively distinct types the letters imply (McCrae & Costa, 1989), and a critical review of the instrument concludes that the four-letter formula does not support the inferences commonly drawn from it (Pittenger, 2005).
A continuous score does not have this problem, because a small wobble produces a small change. This is the biggest single reason to prefer a test that reports where you sit on a scale over one that reports which side of it you fell on.
Why does a result with no evidence behind it still feel accurate?
Bertram Forer gave his students what he described as individual personality sketches. Every student received the same paragraph, assembled from an astrology book, and they rated it as a highly accurate description of themselves, about 4.26 out of 5 as the demonstration is commonly reported (Forer, 1949). Statements that apply to nearly everyone read as uncanny insight when you are told they were written for you.
This is why a format can stay popular for decades without ever being tested. Handwriting analysis is the clearest case. When graphologists' judgements were compared against actual job performance, they did no better than untrained readers working from the same samples, and what predictive value there was came from the content people wrote rather than the shape of the letters (Ben-Shakhar et al., 1986). A meta-analysis of graphological inference reached the same conclusion (Neter & Ben-Shakhar, 1989). People who use it are not being careless; they are meeting the effect Forer demonstrated.
How do the common formats compare?
| Format | Model | Reliability | Validity |
|---|---|---|---|
| Big Five inventories (BFI-2, IPIP-NEO) | Five dimensions, replicated across samples, languages and rater types | Good when each trait has several statements | Good; scores relate to job performance, health behaviour and relationship outcomes |
| HEXACO inventories | Six dimensions, adding Honesty-Humility | Good | Good; the sixth factor covers ground five-factor instruments capture only in part |
| Character strengths surveys | 24 strengths under six virtues | Reasonable | Fair; links to well-being are established, prediction of behaviour less studied |
| Type inventories in the Myers-Briggs tradition | Four dichotomies, each cut at a midpoint | Scales reasonable; the label depends on which side of a midpoint you land | Scales relate to four Big Five traits; the letters add little beyond them |
| Enneagram questionnaires | Nine types | Varies widely between questionnaires | Limited and mixed published evidence |
| Handwriting analysis | No replicated model | Trained raters can agree with each other | No predictive validity found for occupational success |
| Social media personality quizzes | None stated | Not reported | Not studied; written for entertainment |
Reliability and validity are separate columns on purpose. Handwriting analysts can agree with one another and still predict nothing, and a type inventory can use decent scales and still hand you an unstable label.
How can you check a test in five minutes?
- Find the model. Does the page name the Big Five, HEXACO or another published framework, or is the model the product's own invention?
- Count the questions and divide by the number of traits reported. Under five statements per trait, treat the individual scores lightly. A test reporting thirty traits from thirty questions is reporting one question each.
- Look for a result on a scale. A percentile or a band keeps the information. A single label throws most of it away.
- Look for the word norms, or a reference sample, and see whether the test says who is in it.
- Search the test's name alongside reliability and validity. Serious instruments have published papers behind them.
- Read the sample output and ask whether any sentence in it could be false of most people.
How does Wellington approach accuracy?
Wellington is a science-based personality assessment from Therabot Labs LLC. It measures personality as continuous traits, built on the Big Five and HEXACO models, and writes the results back as a personal report rather than a type.
Concretely: the free Snapshot asks 60 statements across six dimensions (Extraversion, Agreeableness, Conscientiousness, Emotional Stability, Openness to Experience, Honesty-Humility), several for each, so no dimension rests on one or two answers. The models are the Big Five and HEXACO rather than a structure of our own. Every trait is reported as a percentile against the reference sample of adults, in one of three bands, Quiet below the 30th percentile, Balanced from the 30th to the 70th, High above. The bands and the Portrait names exist to make the result easier to talk about; the percentile is the measurement, and it is always shown alongside. Those norms are provisional while the norming sample is completed, and we say so on the report and on the methodology page rather than presenting them as settled. The first ten questions are answered before anything is asked of you; after question ten you sign in with an email so your answers are saved and your results can be shown. There is no card.
The membership adds 336 more statements to reach 90 traits, which is what it takes to measure facets with enough questions each to be worth reading. We have not yet published our own reliability figures, and we would rather say that than quote numbers we have not earned; they will go on the methodology page when the norming sample is complete. Wellington is a wellness tool for reflection, not a medical or psychological diagnosis, and it does not replace care from a qualified professional. It is not designed for hiring or selection, and you can export or delete your data at any time.
Questions people ask
- What is the most accurate personality test?
- Among widely available tests, the best evidence sits with Big Five and HEXACO inventories that publish their development and reliability, such as the BFI-2, the IPIP-NEO and the HEXACO-PI-R. They are accurate in the sense that matters: consistent on retest, and related to outcomes in independent samples. No personality test is precise about a single person on a single day.
- How accurate is the Big Five?
- The model itself is the most replicated in personality psychology, recovered independently from trait adjectives, questionnaires and ratings by other people. A well-built inventory places you close to the same position months later, and the trait scores relate to outcomes at work, in health and in relationships. The margin of error is real and a good report states it.
- Is 16Personalities accurate?
- Its makers describe the model as their own, drawing on trait theory and Big Five terminology, and the questionnaire scores continuous scales before reporting a type. The scales are not arbitrary. The weakness is the cut: scales in this format behave as continuous traits rather than as real dichotomies, so people near the middle can receive a different type on a retake. The scores are more informative than the letters.
- Can a ten-question personality test be accurate?
- It can give a rough read on broad dimensions, which is what very short scales were built for, and their authors are explicit that precision is the price. Two statements per trait leave enough noise to move you across a band boundary. Use a very short test to get oriented, and a longer one before you conclude anything about a particular trait.
Sources
Peer-reviewed sources for the claims above. Wellington's own reliability figures will be published once the norming sample is complete.
- Nunnally, J. C., & Bernstein, I. H. (1994). Psychometric theory (3rd ed.). McGraw-Hill.
- Goldberg, L. R. (1993). The structure of phenotypic personality traits. American Psychologist, 48(1), 26–34.
- Ashton, M. C., & Lee, K. (2007). Empirical, theoretical, and practical advantages of the HEXACO model of personality structure. Personality and Social Psychology Review, 11(2), 150–166.
- Credé, M., Harms, P., Niehorster, S., & Gaye-Valentine, A. (2012). An evaluation of the consequences of using short measures of the Big Five personality traits. Journal of Personality and Social Psychology, 102(4), 874–888.
- Gosling, S. D., Rentfrow, P. J., & Swann, W. B. (2003). A very brief measure of the Big-Five personality domains. Journal of Research in Personality, 37(6), 504–528.
- Vazire, S. (2010). Who knows what about a person? The self-other knowledge asymmetry (SOKA) model. Journal of Personality and Social Psychology, 98(2), 281–300.
- Connelly, B. S., & Ones, D. S. (2010). An other perspective on personality: Meta-analytic integration of observers' accuracy and predictive validity. Psychological Bulletin, 136(6), 1092–1122.
- Pittenger, D. J. (2005). Cautionary comments regarding the Myers-Briggs Type Indicator. Consulting Psychology Journal: Practice and Research, 57(3), 210–221.
- McCrae, R. R., & Costa, P. T. (1989). Reinterpreting the Myers-Briggs Type Indicator from the perspective of the five-factor model of personality. Journal of Personality, 57(1), 17–40.
- Forer, B. R. (1949). The fallacy of personal validation: A classroom demonstration of gullibility. Journal of Abnormal and Social Psychology, 44(1), 118–123.
- Ben-Shakhar, G., Bar-Hillel, M., Bilu, Y., Ben-Abba, E., & Flug, A. (1986). Can graphology predict occupational success? Two empirical studies and some methodological ruminations. Journal of Applied Psychology, 71(4), 645–653.
- Neter, E., & Ben-Shakhar, G. (1989). The predictive validity of graphological inferences: A meta-analytic approach. Personality and Individual Differences, 10(7), 737–745.
- Roberts, B. W., Kuncel, N. R., Shiner, R., Caspi, A., & Goldberg, L. R. (2007). The power of personality: The comparative validity of personality traits, socioeconomic status, and cognitive ability for predicting important life outcomes. Perspectives on Psychological Science, 2(4), 313–345.
- Soto, C. J. (2019). How replicable are links between personality traits and consequential life outcomes? The Life Outcomes of Personality Replication project. Psychological Science, 30(5), 711–727.
Read next
See your own traits, not a type.
The free Snapshot takes about seven minutes and gives you your personality card and five dimensions. No credit card, no type.