psychometrics

Are Personality Tests Reliable? What Test-Retest Studies Reveal

Personality tests vary widely in scientific accuracy. The Big Five model (also called the Five Factor Model or OCEAN model) has the strongest research backing, with test-retest reliability coefficients typically between .80 and .90, meaning people get consistent results over time. The MBTI (Myers-Briggs Type Indicator), by contrast, shows much weaker reliability — research suggests that 39% to 76% of people receive a different four-letter type when retaking the test after just five weeks. The key difference comes down to how each test is constructed: the Big Five measures personality traits on a continuous spectrum, while the MBTI forces people into binary categories that don’t reflect how personality actually works. If you want a personality assessment that holds up under scientific scrutiny, trait-based models like the Big Five are the clear winner.

What Makes a Personality Test “Accurate”? Two Concepts You Need to Know

When psychologists talk about whether a personality test is accurate, they look at two distinct properties: reliability and validity. Reliability means consistency — if you take the same test twice, you should get roughly the same result. Validity means the test actually measures what it claims to measure — not just something that sounds similar. A test can be reliable without being valid (imagine a ruler that consistently measures everything as two inches too long), and a valid test that isn’t reliable produces results too noisy to trust.

Reliability is typically measured using a statistic called a correlation coefficient, which ranges from 0 to 1. Higher numbers mean more consistency. Researchers evaluate this through test-retest studies: the same group of people takes the test twice, with days or weeks in between, and the researchers compare the two sets of scores. A well-designed personality test should produce correlation coefficients above .70 for short intervals and above .60 for longer periods spanning months or years.

Validity comes in several forms, but the most important for personality tests is construct validity — whether the test actually captures the psychological trait it’s supposed to measure. Researchers establish construct validity through factor analysis, a statistical method that checks whether test items cluster together the way the theory predicts. They also look at criterion validity: does the test score predict real-world outcomes, like job performance or relationship satisfaction?

Big Five Personality Test: Why Researchers Trust It

The Big Five model measures five broad personality dimensions — Openness, Conscientiousness, Extraversion, Agreeableness, and Neuroticism (often remembered as OCEAN). Each trait represents a spectrum, meaning everyone falls somewhere along a continuum rather than being pigeonholed into a category. The Big Five is considered the gold standard in personality research because its factor structure has been replicated across cultures, languages, and age groups.

A meta-analysis published by Timo Gnambs in 2014 analyzed 682 test-retest correlations from 74 independent samples (total N = 14,923) and found a median dependability coefficient of .816 across all five traits. The NEO Personality Inventory (NEO PI-R), one of the most widely used Big Five instruments, reports test-retest correlations ranging from .85 to .92 over several weeks. Even over a six-year span, researchers Costa and McCrae found correlations between .68 and .83 for the five domain scores — remarkable stability for a psychological measure.

The Big Five also demonstrates strong criterion validity. According to the landmark meta-analysis by Barrick and Mount published in Personnel Psychology, Conscientiousness predicts job performance across virtually all occupational groups. Neuroticism is the strongest personality predictor of mental health difficulties, while Agreeableness predicts relationship quality. These findings have been replicated in dozens of studies, making the Big Five the model that other personality tests are measured against.

MBTI Reliability Issues: Why Your Type Keeps Changing

The MBTI sorts people into 16 personality types based on four binary dimensions: Introversion vs. Extraversion, Sensing vs. Intuition, Thinking vs. Feeling, and Judging vs. Perceiving. The problem is that personality traits don’t distribute as either/or categories — they follow a normal distribution (a bell curve), with most people clustering near the middle. When a test forces a continuous trait into a binary split, people near the middle flip-flop between categories on retest, producing inconsistent results.

Research by Jim Pittenger, published in 2005, found that 39% to 76% of test-takers receive a different four-letter type code when retaking the MBTI after five weeks. A meta-analysis by Capraro and Capraro found per-dimension test-retest correlations of approximately .50 to .60 — meaningfully lower than the Big Five’s .80+ range. Even the MBTI’s own publisher acknowledges these limitations: the Form M Manual Supplement reports per-dimension correlations ranging from .53 to .93 depending on the scale and time interval, with the lower end falling well below accepted psychometric standards.

The MBTI also faces construct validity challenges. Factor analysis — the statistical technique used to verify that test items group as expected — does not consistently reproduce the MBTI’s four independent dimensions. Instead, studies tend to find that the eight MBTI poles map onto Big Five traits, suggesting the MBTI is measuring something real but capturing it less precisely than the Five Factor Model does.

Does This Mean the MBTI Is Useless?

Not necessarily. The MBTI’s popularity stems from its accessibility — the 16-type framework gives people a memorable language for thinking about themselves and others. Many people find it genuinely useful for self-reflection and team conversations. The issue isn’t that the MBTI captures nothing real; it’s that it oversimplifies. A common question people ask is: “Why do I get different MBTI results every time?” The answer is that anyone scoring near the midpoint on any dimension will be classified differently based on small day-to-day variations in mood or self-perception.

For casual self-discovery, the MBTI can serve as a starting point. For decisions that matter — hiring, clinical assessment, research — personality psychologists overwhelmingly recommend the Big Five. The MBTI’s publisher explicitly states it should not be used for hiring or selection decisions.

Self-Report Bias: The Challenge Every Personality Test Faces

Both the Big Five and MBTI rely on self-report — asking people to rate statements about themselves. This introduces a fundamental challenge: people don’t always see themselves accurately. Research on self-other agreement shows that people’s self-ratings correlate only moderately (around .40 to .60) with how others rate them, meaning there’s a meaningful gap between self-perception and observable behavior.

Several factors distort self-report accuracy. Social desirability bias leads people to present themselves favorably — for example, rating themselves higher on Conscientiousness than their behavior warrants. Reference group effects mean that people compare themselves to those around them: someone might rate themselves as highly extraverted because their friends are introverted, even if they’d score average in a broader sample. Mood and context also matter: taking a personality test after a difficult week can temporarily lower scores on Emotional Stability (the inverse of Neuroticism).

Test designers address these issues in several ways. Many instruments include “lie scales” or social desirability checks — items designed to catch inconsistent or overly favorable responding. The Big Five’s use of spectrum scoring (rather than binary cutoffs) makes it more resilient to small biases, because a slight overestimate still places you in roughly the right range. Some researchers also use informant reports — asking friends, family, or colleagues to rate the person — which can provide a more accurate picture than self-report alone.

How to Choose a Personality Test You Can Trust

If you’re looking for a personality assessment that balances scientific rigor with practical usefulness, here are some guidelines:

  • Look for spectrum scoring. Tests that place you on a continuum (like the Big Five) are more reliable and valid than those that assign you to a fixed category (like the MBTI).
  • Check the test-retest reliability. A well-designed test should report correlations above .70 for short intervals. If the test publisher doesn’t disclose reliability data, that’s a red flag.
  • Prefer transparency over mystique. Legitimate personality tests explain their methodology and cite research. Tests that promise to reveal “hidden” aspects of your personality often lack scientific grounding.
  • Use results as a starting point, not a verdict. Personality is complex and context-dependent. A test score is a snapshot, not a permanent label.

If you want to discover your own personality profile, tools like personalitree.com offer free Big Five and 16-type assessments that take about 10 minutes and provide results based on established personality frameworks. Taking both the Big Five and a 16-type test can give you complementary perspectives — the Big Five for scientific precision, and the 16 types for an intuitive shorthand.

What Personality Tests Can and Cannot Tell You

Personality tests are tools for self-understanding, not crystal balls. A well-validated test like the Big Five can tell you where you fall relative to others on five major trait dimensions, and research links those scores to real-world outcomes — from career fit to relationship patterns to mental health tendencies. What tests cannot do is predict your future, determine your potential, or box you into a fixed identity.

Research from Wright and colleagues, published in Communications Psychology in 2026, analyzed data from 167,000 participants across eight countries and found that Big Five traits change meaningfully over the lifespan — people become more Conscientious, more Agreeable, and more Emotionally Stable with age. This means your test results today aren’t permanent. Personality is neither fixed nor infinitely malleable — it’s a set of tendencies that evolve with experience, environment, and intentional growth.

Frequently Asked Questions

Why do I get different personality test results each time I take the test?

This is common with tests that use binary categories, like the MBTI. If you score near the middle on any dimension, small variations in your mood or self-perception can push you into a different category. Spectrum-based tests like the Big Five are more stable because they measure traits on a continuum — you might shift slightly, but you won’t suddenly jump from one type to another. Research shows that 39% to 76% of people get a different MBTI type on retest within five weeks.

Is the Big Five personality test more accurate than the MBTI?

Yes, according to the consensus of personality researchers. The Big Five has stronger test-retest reliability (typically .80 to .90 versus the MBTI’s .50 to .60 per dimension), better construct validity confirmed through factor analysis, and demonstrated criterion validity — its traits predict real-world outcomes like job performance and mental health. The MBTI can still be useful for self-reflection, but it’s less precise and less stable.

Can personality tests be wrong about me?

Yes. All self-report personality tests are subject to biases like social desirability (answering in a flattering way), reference group effects (comparing yourself to the people around you), and temporary mood states. Your results reflect your self-perception at the time of testing, which may differ from how others see you or how you behave in different contexts. Using results as a general guide rather than a definitive label helps account for this uncertainty.

What is a good reliability score for a personality test?

Psychologists generally consider test-retest reliability above .70 to be acceptable and above .80 to be good. The NEO PI-R (a Big Five instrument) reports correlations of .85 to .92 over several weeks, while shorter measures like the BFI-2 report around .80. Tests with reliability below .60 — as some MBTI dimensions show — produce results that are too inconsistent for confident interpretation.

Should I use a personality test for hiring decisions?

Only if you use a well-validated instrument like the Big Five, and even then, with caution. The MBTI’s own publisher states it should not be used for selection or hiring. Research shows that Conscientiousness (a Big Five trait) modestly predicts job performance across occupations, but no personality test should be the sole basis for a hiring decision. Personality data works best as one input among many, combined with interviews, skills assessments, and work samples.

Are Personality Tests Reliable? What Test-Retest Studies Reveal Read More »

Why Some Personality Tests Work and Others Are Pseudoscience

Every day, millions of people take personality tests online. Some are looking for career guidance, others want to understand their relationships better, and many are simply curious about what a test might reveal. But behind the colorful result pages, type descriptions, and percentage breakdowns lies a rigorous scientific discipline called psychometrics — the study of psychological measurement. Understanding how personality tests are actually built, validated, and scored can help you tell the difference between a test grounded in decades of research and one that is essentially a sophisticated horoscope.

The personality testing industry has grown dramatically over the past decade. The global psychometric testing market was valued at several billion dollars and continues to expand as organizations integrate personality assessments into hiring, team development, and leadership training. Yet the quality gap between the best and worst tests is enormous. A well-constructed Big Five inventory, developed through years of factor analysis and validated across diverse populations, shares almost nothing in common with a ten-question quiz designed to generate social media engagement. Knowing what separates them matters.

How Personality Tests Are Built: The Item Construction Process

Building a scientifically valid personality test is not a matter of brainstorming questions that sound insightful. The process follows a structured methodology that can take years from initial concept to published instrument.

The first stage is construct definition. Before writing a single question, test developers must clearly define what they are trying to measure. For the Big Five model, this meant decades of lexical research — analyzing thousands of personality-descriptive words across multiple languages and using factor analysis to identify the underlying dimensions that consistently emerged. Researchers like Lewis Goldberg, Paul Costa, and Robert McCrae demonstrated that personality descriptions cluster around five broad factors regardless of culture, language, or measurement method. This cross-cultural replication is one of the strongest arguments for the Big Five’s validity.

Once the construct is defined, item writing begins. Test developers generate a large pool of potential questions — often hundreds — designed to tap into the target trait. Good items are clear, specific, and behaviorally anchored. Rather than asking “Are you creative?” which invites vague self-assessment, a better item might ask “How often do you generate unusual ideas?” with a frequency-based response scale. The wording must avoid social desirability bias, double-barreled phrasing, and cultural references that would not translate across populations.

The initial item pool then undergoes pilot testing with a representative sample. Statistical analyses — including item-total correlations, difficulty indices, and differential item functioning tests — identify which items perform well and which need revision or removal. Items that do not correlate with the overall scale, that show bias across demographic groups, or that fail to discriminate between high and low scorers on the trait are eliminated. This iterative process can reduce an initial pool of 200 items to a final set of 40 or 50 that measure the construct cleanly.

Reliability: Can the Test Produce Consistent Results?

Reliability refers to consistency. If you take a personality test on Monday and again on Friday, you should get roughly the same results — assuming nothing major happened in between. In psychometrics, reliability is quantified through several methods, each addressing a different aspect of consistency.

Internal consistency, measured by Cronbach’s alpha, assesses whether all items on a given scale are measuring the same underlying construct. A Cronbach’s alpha above 0.70 is generally considered acceptable for research purposes; above 0.80 is good; and above 0.90 is excellent. The official MBTI assessment reports Cronbach’s alpha values around 0.90 for its scales, while well-constructed Big Five inventories routinely achieve similar or higher values. A test with low internal consistency is essentially measuring noise alongside signal — you cannot trust its individual scale scores because the items do not cohere.

Test-retest reliability measures stability over time. A person’s score on Extraversion should not change dramatically from one week to the next. Research on Big Five inventories typically finds test-retest correlations in the 0.80-0.90 range over periods of weeks to months. The MBTI shows test-retest reliability around 0.81-0.86 over one to six weeks, though some studies have found lower stability for certain dimensions, particularly the Thinking-Feeling and Judging-Perceiving scales. When a test shows poor test-retest reliability, it means the results are heavily influenced by momentary mood, testing context, or random error rather than stable personality traits.

Inter-rater reliability is less commonly reported for self-report personality tests but becomes relevant in observer-report versions. When a test asks someone who knows you well to rate your personality, their ratings should correlate meaningfully with your self-ratings. Research consistently finds moderate to strong self-other agreement on Big Five traits, with correlations typically in the 0.40-0.60 range, which is substantial given that different raters have access to different behavioral information.

Validity: Does the Test Measure What It Claims to Measure?

Reliability is necessary but not sufficient. A test can produce perfectly consistent results that are consistently wrong. Validity addresses whether the test actually measures the construct it claims to measure.

Content validity asks whether the test items adequately cover the full breadth of the construct. A conscientiousness scale that only asks about punctuality misses the broader dimensions of the trait — organization, diligence, achievement striving, and self-discipline. Test developers establish content validity through expert review panels and systematic mapping of items to the construct’s theoretical components.

Criterion validity — often divided into concurrent and predictive validity — examines whether test scores correlate with real-world outcomes. The Big Five shows impressive criterion validity across multiple domains. Conscientiousness predicts job performance across virtually all occupations, with meta-analytic correlations in the 0.20-0.30 range. Neuroticism predicts vulnerability to anxiety and depression. Extraversion predicts leadership emergence and sales performance. These correlations may seem modest, but in psychological research, where outcomes are determined by many factors, they represent meaningful predictive power.

Construct validity is the broadest form of validity evidence — it asks whether the pattern of relationships between the test and other measures matches theoretical expectations. A valid Extraversion scale should correlate positively with measures of social engagement and positive affect, correlate negatively with social anxiety, and show near-zero correlations with unrelated constructs like numerical ability. The Big Five has accumulated overwhelming construct validity evidence over decades of research. The MBTI, by contrast, has faced more criticism in this area, particularly regarding its binary type categories and the theoretical independence of its four dimensions.

The Big Five vs. 16 Personalities: A Tale of Two Frameworks

The scientific standing of the Big Five and the 16 Personalities model differs significantly, and understanding why illuminates what makes a personality test credible.

The Big Five emerged from the lexical approach — the observation that the most important personality differences between people become encoded in language over time. By analyzing personality-descriptive adjectives across languages and applying factor analysis, researchers repeatedly found five broad dimensions. The model is descriptive (it summarizes what traits exist) rather than theoretical (it does not claim to explain why they exist), which grounds it in empirical observation. The Big Five has been replicated across cultures, age groups, and measurement methods, and it predicts a wide range of life outcomes including academic achievement, job performance, relationship satisfaction, and even longevity.

The 16 Personalities model, rooted in Carl Jung’s theory of psychological types and operationalized by Katharine Cook Briggs and Isabel Briggs Myers, takes a different approach. It sorts people into 16 discrete categories based on four dichotomies: Extraversion-Introversion, Sensing-Intuition, Thinking-Feeling, and Judging-Perceiving. The modern 16Personalities website adds a fifth dimension — Assertive-Turbulent, mapping onto the Big Five’s Neuroticism — in what is called the NERIS model, bridging the two frameworks.

The MBTI’s scientific criticisms are well-documented. The binary categories impose cutoffs on continuous distributions, meaning two people with nearly identical scores on a dimension can be classified into opposite types. The test-retest reliability of the type categories is lower than that of dimensional scores, with studies finding that 39-76% of test-takers receive a different type classification upon retesting. And the theoretical independence of the four dimensions has not been consistently supported by factor analysis. Despite these limitations, the MBTI remains enormously popular because it provides accessible language, positive framing of all types, and a sense of identity that dimensional models do not offer as intuitively.

If you want to explore your own personality type, platforms like personalitree.com offer free assessments that cover both frameworks — the Big Five for scientific rigor and dimensional nuance, and the 16-type model for accessible self-reflection and discussion. Having both perspectives gives you a more complete understanding than either framework alone.

What Makes a Test Worth Taking: A Practical Checklist

Given the wide variation in test quality, how can a non-specialist evaluate whether a personality test is worth the time it takes to complete? Several indicators separate scientifically grounded assessments from entertainment.

First, look for transparency about the test’s development. A credible test will name the specific model it uses (not a vague “personality type” framework), cite the research behind it, and report its psychometric properties — reliability coefficients, validity evidence, and the characteristics of its norming sample. If a test website provides no information about how the test was developed or validated, proceed with skepticism.

Second, examine the item quality. Scientifically constructed items ask about specific, observable behaviors rather than abstract self-assessments. They avoid leading language, extreme wording, and items where one response is clearly more socially desirable. A test with vague, repetitive, or poorly translated items is unlikely to produce meaningful results.

Third, consider the response format. The most reliable personality tests use Likert-type scales — typically five or seven points from “strongly disagree” to “strongly agree” — rather than binary yes/no or forced-choice formats. Dimensional response scales capture more information and better reflect the continuous nature of personality traits.

Fourth, check the length. While there is no magic number, a personality test with fewer than 30-40 items is unlikely to measure multiple traits with adequate reliability. The full NEO-PI-R, one of the most respected Big Five instruments, contains 240 items. Shorter scales exist and can be useful, but extreme brevity comes at the cost of precision.

Fifth, be wary of overly specific predictions. A legitimate personality test describes broad patterns and tendencies, not specific life outcomes. Any test that claims to predict your ideal career with certainty, identify your perfect romantic partner, or reveal hidden truths about your destiny is selling something other than psychological science.

The Limits of Self-Report and What Comes Next

Even the best personality tests face inherent limitations, most notably the self-report problem. When you answer questions about yourself, your responses are filtered through self-perception, which is imperfect. People may lack self-awareness, respond according to how they wish to be rather than how they are, or be influenced by their current mood and recent experiences. Research on self-enhancement bias shows that people tend to rate themselves higher on socially desirable traits like Conscientiousness and Agreeableness and lower on Neuroticism than observer ratings would suggest.

Emerging approaches aim to address these limitations. Observer-report versions of personality inventories ask people who know you well to rate your traits, and the combination of self and observer ratings often provides more predictive power than either alone. Behavioral measures — tracking actual behavior patterns through digital footprints, language analysis, or structured observation — offer another path forward, though these methods raise significant privacy concerns. Some researchers are exploring implicit measures that assess automatic associations rather than conscious self-descriptions, though the predictive validity of these approaches remains debated.

For most people, the practical takeaway is straightforward: personality tests are tools, not oracles. They provide structured information that can spark useful self-reflection, highlight patterns you might not have noticed, and offer a vocabulary for discussing differences with others. A well-validated test from a credible source — such as those based on the Big Five model available through websites like personalitree.com — can be a valuable starting point for self-understanding. The test does not define you; it describes tendencies that you can choose to work with, work around, or work on.

Why Some Personality Tests Work and Others Are Pseudoscience Read More »