Skip to content
Seven minutes. A new way to see yourself.Get your free results →
Wellington

Glossary

What is Cronbach's alpha, and what does it tell you about a test?

Cronbach's alpha is a number, usually between 0 and 1, that says how consistently the statements meant to measure one trait move together within a single sitting. A high alpha means people who agree with one statement on the scale tend to agree with the others too. It is one form of reliability, not a measure of whether the scale is measuring the right thing (Nunnally & Bernstein, 1994).

Last updated September 16, 2026.

What is Cronbach's alpha, in plain words?

Picture a scale for Conscientiousness built from eight statements about tidiness, follow-through and planning ahead. If the scale is working, a person who agrees strongly with one statement should tend to agree with the others too, because all eight are meant to tap the same underlying trait. Cronbach's alpha is a single number that captures how true that is across a group of respondents, for one scale, answered once.

Mechanically, it is built from the correlations between every pair of items on the scale, folded into one figure that behaves like a proportion: the closer to 1, the more the items move together; the closer to 0, the more they behave like unrelated statements sharing a page. It is one form of internal consistency, alongside test-retest reliability and the other kinds covered on are personality tests reliable. Reverse-keyed statements, where agreeing points toward the low end of the trait, are flipped before this calculation runs, for the same reason they are flipped in the trait score itself; see reverse-keyed items.

Why does adding more items raise alpha?

Every single statement carries noise: an odd mood on the day, a word read slightly differently than intended, a moment of second-guessing. Averaging several statements about the same trait cancels out much of that noise, because one statement's quirks rarely line up with another's. Classical test theory treats this as close to a mechanical fact: holding the average correlation between items constant, alpha rises as more items are added to the scale, which is why careful inventories tend to be longer than they look as though they need to be (Nunnally & Bernstein, 1994).

That relationship cuts both ways. A short scale is mathematically disadvantaged before a single respondent has answered a single item. It also means alpha can be nudged upward simply by adding items that restate each other in slightly different words, which raises the number without adding what a good test actually wants: coverage of the trait's different corners. A high alpha is evidence the items are moving together, not proof the scale was well built.

The practical cost of going short has been measured rather than assumed. Using data from working adults and college students, one evaluation found that very abbreviated personality measures, single-item measures especially, led researchers to underestimate how much personality traits actually predicted real outcomes, and to overstate the importance of other constructs measured alongside them; scales only slightly longer improved the picture substantially, for little extra time from the person answering (Credé et al., 2012).

What counts as a good Cronbach's alpha?

There is no single number that settles the question, and no citation that fixes one. Textbooks on measurement discuss thresholds with real qualifications attached, because the right bar depends on what the score is used for: a screening decision about one person calls for more consistency than a group comparison in a study, and a scale with very few items will struggle to clear a bar written with longer scales in mind. With that caveat in front of it, alpha around .7 or higher is often described as a common rule of thumb for acceptable internal consistency, and .8 or above as comfortable. Treat this as convention rather than a law: a scale can sit below .7 and still be useful for its purpose, and one well above .9 can mean the items are so similar to each other that little is gained by asking all of them rather than a handful.

What do real personality tests report for alpha?

Alpha is easiest to make concrete against instruments that publish it. The IPIP-NEO-120, a public-domain measure built to cover all 30 facets of the five-factor model with just four items per facet, showed acceptable reliability across all 30 of those facet scales in two large internet samples, evidence that a fairly short, open item pool can still cover the ground a facet needs (Johnson, 2014).

The HEXACO Personality Inventory, built around 24 facet-level scales, showed high internal consistency across those facets in a validation sample of more than 400 respondents (Lee & Ashton, 2004). The BFI-2, a hierarchical measure with five domains and 15 facets nested beneath them, was built and validated so its facet scores add predictive power over the domain scores alone (Soto & John, 2017); its 60 items work out at 12 per domain and 4 per facet, and its authors make it available at no cost for research and educational use. None of these figures transfers to a different item bank; alpha belongs to one set of items, answered by one sample, which is why a trustworthy report names the scale a number belongs to.

What kinds of reliability does alpha leave out?

Internal consistency is one kind of reliability among several, and alpha only speaks to the first row
Kind of reliabilityWhat it checksWhere alpha fits
Internal consistency (Cronbach's alpha)Whether the statements on one scale move together, in a single sittingThis is what alpha measures
Split-half reliabilityWhether one half of the items on a scale correlates with the other halfA related single-sitting check; alpha is closer to the average of every possible split
Test-retest reliabilityWhether the same person scores similarly weeks or months apart, with no real change betweenAlpha says nothing about this; a scale can have high internal consistency and still drift on a retake
Self-observer agreementWhether a self-report lines up with what someone who knows the person well would sayA separate question again; it concerns different raters, not different items

A high alpha also says nothing about validity, whether the scale measures what it claims to. A set of items can move together very consistently while all of them miss the trait they were meant to capture, in the same way a bathroom scale can read five pounds heavy every time and still be perfectly consistent. Personality test validity covers what validity evidence looks like, and how the Snapshot was built and scored covers how an item bank is checked before it goes live.

How does Wellington measure this?

Wellington is a science-based personality assessment from Therabot Labs LLC. It measures personality as continuous traits, built on the Big Five and HEXACO models, and writes the results back as a personal report rather than a type.

Concretely: the free Snapshot measures Extraversion, Agreeableness, Conscientiousness, Emotional Stability, Openness to Experience, Honesty-Humility across 60 statements, several per dimension, so a single ambiguous answer does not carry a whole trait score on its own, and so an internal-consistency figure can be calculated for each dimension rather than resting on one item. The first ten questions are answered before anything is asked; after question ten you sign in with an email so your answers are saved and results can be shown, and there is no card. The Wellington Membership extends this to 90 traits, with enough statements per trait to reach facets rather than broad dimensions alone. Every trait is reported as a percentile against the reference sample. The three bands (Quiet below the 30th percentile, Balanced to the 70th, High above) and the Portrait names exist to make the result easier to talk about; the percentile is the measurement, and it is always shown alongside.

We have not yet published our own reliability figures, including alpha for each dimension. We would rather say that plainly than quote a number we have not earned; those figures will appear on the methodology page once the norming sample is complete, and the norms stay provisional until then. Wellington is a wellness tool for reflection, not a medical or psychological diagnosis, and it does not replace care from a qualified professional. It is not designed for hiring, and you can export or delete your data at any time.

Questions people ask

Is a high Cronbach's alpha the same as a good test?
No. Alpha only says the items on a scale moved together for the sample that answered them. A scale can have a high alpha and still miss the trait it claims to measure, cover only a narrow slice of it, or fail to predict anything real. Alpha is a precondition worth checking, not the whole case for a test; that case also needs validity evidence, covered in personality test validity.
Why do some short personality tests still report a decent alpha?
Alpha rises as the average correlation between items rises, not only as items are added, so a short scale built from a few strongly related statements can still clear a common threshold. What that scale usually gives up is coverage: fewer items means fewer angles on the trait (Credé et al., 2012).
Does Cronbach's alpha tell you if a test is consistent over time?
No. Alpha is calculated from one sitting; it says whether the items agreed with each other on that occasion, not whether the same person would score similarly weeks later. That is test-retest reliability, a separate calculation covered in are personality tests reliable, and a test can be strong on one and weaker on the other.
What is considered an acceptable Cronbach's alpha for a personality scale?
There is no fixed cutoff that applies everywhere. Alpha around .7 or higher is often described as a common rule of thumb, with .8 or above seen as comfortable, but the right bar depends on how the score is used and how many items the scale has. Treat any single number as a convention, not a pass or fail line.
Can a scale's alpha be too high?
In a sense, yes. Once alpha climbs well above .9, it often means several items are close restatements of each other rather than covering different corners of the trait, so little is gained by asking all of them instead of a shorter set. A very high alpha is a sign to check for redundancy, not a sign of a better scale.

Sources

Peer-reviewed sources for the claims above. Wellington's own reliability figures will be published once the norming sample is complete.

  1. Nunnally, J. C., & Bernstein, I. H. (1994). Psychometric theory (3rd ed.). McGraw-Hill.
  2. Credé, M., Harms, P., Niehorster, S., & Gaye-Valentine, A. (2012). An evaluation of the consequences of using short measures of the Big Five personality traits. Journal of Personality and Social Psychology, 102(4), 874–888.
  3. Johnson, J. A. (2014). Measuring thirty facets of the Five Factor Model with a 120-item public domain inventory: Development of the IPIP-NEO-120. Journal of Research in Personality, 51, 78–89.
  4. Lee, K., & Ashton, M. C. (2004). Psychometric properties of the HEXACO Personality Inventory. Multivariate Behavioral Research, 39(2), 329–358.
  5. Soto, C. J., & John, O. P. (2017). The next Big Five Inventory (BFI-2): Developing and assessing a hierarchical model with 15 facets to enhance bandwidth, fidelity, and predictive power. Journal of Personality and Social Psychology, 113(1), 117–143.

Read next

See your own traits, not a type.

The free Snapshot takes about seven minutes and gives you your personality card and five dimensions. No credit card, no type.

One trait a week

Not ready to take the test? Read one trait a week.

A real page from the report, one practice attached, every week. It is the easiest way to see whether Wellington reads people the way you think it should.