Methods
How the Wellington Snapshot was built and scored
The free Snapshot asks 60 questions, ten for each of 6 broad dimensions, on a 5-point accuracy scale. Each dimension is scored as the average of its questions, placed against a reference sample as a percentile, and reported in one of three bands with a written page for that band. The Wellington Membership adds 336 questions across 4 chapters to reach 90 traits.
Last updated September 16, 2026.
What does the Snapshot ask?
60 short statements, each rated from "Very inaccurate" to "Very accurate". Ten belong to each of the six dimensions: Extraversion, Agreeableness, Conscientiousness, Emotional Stability, Openness to Experience, Honesty-Humility. The first ten are answered before any email is asked for, so you can see the questions before deciding to continue. It takes about 7 minutes.
The full item bank holds 1,605 statements across 255 traits, so that every trait can be measured with several statements rather than one. Several items per trait is the main defence against a single odd answer moving a score.
The statements themselves are written by our team against the published literature for each construct: the five broad dimensions as the lexical and questionnaire traditions defined them (Goldberg, 1993; Costa & McCrae, 1992), the sixth dimension as the HEXACO model defines it (Ashton & Lee, 2007), and, with the membership, the 24 character strengths as they are classified in the standard handbook (Peterson & Seligman, 2004). The statements are real public-domain items from the International Personality Item Pool and public-domain items written in the same style, curated for a wellness setting; each trait's methods note names the published scale it draws on (Goldberg et al., 2006).
Every statement is about ordinary behaviour, in the present tense, in plain words, and short enough to answer in a few seconds. Items that needed a second reading in testing were rewritten or dropped.
Why does each trait need several questions?
Because one question is mostly noise. The reliability of a scale rises as you add items that tap the same underlying trait, which is the oldest quantitative result in psychometrics and the reason serious inventories are longer than they look as though they need to be (Nunnally & Bernstein, 1994). A single item carries whatever the respondent happened to think that word meant, plus their mood, plus how the sentence was phrased. Ten items average most of that away.
The cost of going short has been measured rather than assumed. Very short Big Five measures were shown to understate how much the traits matter for behaviour and to inflate the apparent importance of newer constructs, single-item measures worst of all, while even slightly longer scales improved validity substantially at little cost in a respondent's time (Credé et al., 2012). Very brief instruments have their place: the authors of a widely used ten-item measure said plainly that it is inferior to a standard multi-item inventory and should be chosen only when time is severely limited and the alternative is no measure at all (Gosling et al., 2003).
Ten items per dimension is our compromise between precision and the 7 minutes most people will actually give a free assessment. It is enough for a domain score and not enough for facets, which is why the facets sit in the Portrait rather than in the Snapshot.
How is a score made?
- Reverse-keyed statements are flipped, so agreeing with a statement phrased against the trait counts the right way.
- The trait score is the average of its statements on the five-point scale.
- The score is compared with a reference sample of adults to give a standard score, then a percentile: the share of the reference sample scoring below you.
- Percentiles below the 30th are reported as the Quiet band, the 30th to the 70th as Balanced, and above the 70th as High. The bands are wide on purpose; a percentile near a boundary can fall either side on a retake, and the report says so.
| Band | Percentile range | How to read it |
|---|---|---|
| Quiet | Below the 30th | Lower than most adults in the reference sample on this trait, with the advantages that position carries as well as the costs |
| Balanced | 30th to 70th | Near the middle, which is where most people are; the written page describes what a mixed position looks like rather than splitting the difference |
| High | Above the 70th | Higher than most adults in the reference sample, again with both sides described |
Our reference norms are provisional while the norming sample is completed, and the report's methods section states this. The bands and the Portrait names exist to make the result easier to talk about; the percentile is the measurement, and it is always shown alongside. Scores near the middle of the scale should be read as balanced, not as an absence of the trait.
A percentile is a comparison, not a quantity. It says where you sit relative to a particular group of adults, so changing the group would change the number without anything changing about you. How to read personality test scores works through what a band does and does not license you to conclude.
Where does the writing come from?
In the written report, every one of the 255 traits has three pages, one per band, 765 in all, written in advance by our writing and psychology team and grounded in the research on that trait. Your answers choose which pages you read. The report then adds sections built from your own pattern: an overview of each domain, a constellation of your highest strengths, and a growth plan chosen from your lowest bands, each with a small practice.
Writing the pages in advance is a deliberate constraint. It means the text you read has been checked before anyone reads it, that the page for a low band is written with the same care as the page for a high one, and that nothing in your report was improvised around your answers.
What are the limits of a self-report?
Everything above assumes you are a reasonable judge of yourself, and you are, unevenly. People know their own inner states better than observers do, which makes self-report the stronger source for traits like Emotional Stability that are largely felt rather than seen. Others do better on the traits that carry an evaluation, such as intellect, where the person's own view is the one most likely to be flattering, while on highly visible traits such as extraversion self and others do about equally well (Vazire, 2010).
This is not an argument for distrusting your own answers. It is an argument for knowing which parts of the report to hold loosely. Observer ratings predict behaviour about as well as self-ratings, and better for some outcomes such as academic achievement and job performance, so the two sources are best read together rather than one instead of the other (Connelly & Ones, 2010). If a band surprises you, the quickest test is to show the page to someone who has known you for years.
- Mood. A hard week can move a score a few points. Retaking after a normal one is fair, not cheating.
- Self-presentation. Answering as you would like to be rather than as you are shows up as an unusually flattering profile across every dimension at once.
- Reference group. You rate yourself against the people you happen to know, which is one more reason percentiles against a stated sample are more useful than raw averages.
How does the Wellington Membership extend it?
The Portrait is 336 further statements in 4 chapters, about 36 minutes in total, which you can take on different days. It brings the report to 90 traits, including every facet beneath the six dimensions and all 24 VIA character strengths. The Wellington Membership continues from there with a short module each week for 104 weeks, adding traits from emotional functioning, relationships, stress and resilience, meaning and values, lifestyle and self-concept until all 255 are measured, then retesting the earliest ones so you can see change.
| Questions | Traits measured | Cost | |
|---|---|---|---|
| Snapshot | 60 | 6 broad dimensions | Free |
| Wellington Membership | 336 more, then a short module each week | 90 traits including every Big Five facet and 24 character strengths, growing toward all 255 across eight domains with retests in year two | $4.99 a month, cancel any time |
The facets are the reason the Portrait exists. Two people at the same percentile on Conscientiousness can differ completely on whether that score is orderliness or industriousness, and a domain score cannot tell them apart. The facets page explains what sits beneath each dimension, and the glossary defines the dimensions themselves.
What does it not do?
Wellington is a wellness tool for reflection, not a medical or psychological diagnosis, and it does not replace care from a qualified professional. It does not screen for, detect or rule out any condition. It is not designed for hiring or selection. You can export or delete your data any time from your dashboard; the privacy policy has the full statement on how your data is handled.
It also does not claim more precision than it has. We have not yet published reliability coefficients for our own scales, because the norming sample is not complete, and we would rather say that than quote a number we have not earned. When those figures exist they will appear here. How accurate are personality tests sets out the standard we expect to be held to, and the evidence page covers the models underneath.
Questions people ask
- How reliable is the Snapshot?
- Each dimension is measured with ten statements drawn from a bank built on the Big Five and HEXACO literature, which is the standard defence against noise. We will publish our own reliability figures once the norming sample is complete rather than quote numbers we have not yet earned.
- Why six dimensions and not five?
- Because lexical studies run from scratch in several languages produce six broad factors rather than five. The sixth, Honesty-Humility, separates sincerity, fairness and modesty out from Agreeableness, and it carries information the five-factor model folds in and loses (Ashton & Lee, 2007). Measuring it costs ten more questions in the Snapshot and tells you something the Big Five alone would not.
- Can I retake it?
- Yes, from your dashboard, whenever you like. The newest completed sitting is the one your report is built from, and older sittings are kept so you can compare them. A retake is always a fresh set of answers rather than an edit of the old ones, which is what makes the comparison meaningful. A hard week can move a score a few points, so retaking after an ordinary one is fair.
- Can I see the questions before I start?
- Yes. The first ten of the 60 questions are answered before anything is asked of you, so you can read the statements and decide whether the style suits you. After question ten you sign in with an email address, so your answers are saved and your results can be shown to you. No card is involved at any point in the free Snapshot.
- Why ten questions per dimension rather than two?
- Because precision rises with the number of items measuring the same trait (Nunnally & Bernstein, 1994), and very short forms have been shown to understate how much a trait matters for behaviour (Credé et al., 2012). Ten is our compromise between accuracy and the time people will give a free assessment. The facets need more, which is why they sit in the Portrait rather than the Snapshot.
- What are the percentiles compared against?
- A reference sample of adults. Your percentile is the share of that sample scoring below you on the trait, which is why the same answers would produce a different number against a different group. Our norms are provisional while the norming sample is completed, and the report says so rather than hiding it in a footnote.
Sources
Peer-reviewed sources for the claims above. Wellington's own reliability figures will be published once the norming sample is complete.
- Goldberg, L. R. (1993). The structure of phenotypic personality traits. American Psychologist, 48(1), 26–34.
- Costa, P. T., & McCrae, R. R. (1992). Revised NEO Personality Inventory (NEO PI-R) and NEO Five-Factor Inventory (NEO-FFI) professional manual. Psychological Assessment Resources.
- Ashton, M. C., & Lee, K. (2007). Empirical, theoretical, and practical advantages of the HEXACO model of personality structure. Personality and Social Psychology Review, 11(2), 150–166.
- Peterson, C., & Seligman, M. E. P. (2004). Character Strengths and Virtues: A Handbook and Classification. Oxford University Press and the American Psychological Association.
- Goldberg, L. R., Johnson, J. A., Eber, H. W., Hogan, R., Ashton, M. C., Cloninger, C. R., & Gough, H. G. (2006). The international personality item pool and the future of public-domain personality measures. Journal of Research in Personality, 40(1), 84–96.
- Nunnally, J. C., & Bernstein, I. H. (1994). Psychometric theory (3rd ed.). McGraw-Hill.
- Credé, M., Harms, P., Niehorster, S., & Gaye-Valentine, A. (2012). An evaluation of the consequences of using short measures of the Big Five personality traits. Journal of Personality and Social Psychology, 102(4), 874–888.
- Gosling, S. D., Rentfrow, P. J., & Swann, W. B. (2003). A very brief measure of the Big-Five personality domains. Journal of Research in Personality, 37(6), 504–528.
- Vazire, S. (2010). Who knows what about a person? The self-other knowledge asymmetry (SOKA) model. Journal of Personality and Social Psychology, 98(2), 281–300.
- Connelly, B. S., & Ones, D. S. (2010). An other perspective on personality: Meta-analytic integration of observers' accuracy and predictive validity. Psychological Bulletin, 136(6), 1092–1122.
Read next
See your own traits, not a type.
The free Snapshot takes about seven minutes and gives you your personality card and five dimensions. No credit card, no type.