What Is Reliability? | Definition, Types & Examples

Reliability in research refers to how consistently a research method measures something. If you use the same method on the same group under the same conditions, you should get similar results each time. If the results vary significantly, the measurement may be unreliable or affected by bias or random error.

Researchers use reliability to evaluate whether a test, questionnaire, interview, observation, or other measurement tool produces dependable results. A reliable method increases confidence that differences in results reflect real differences, not inconsistencies in the measurement itself.

There are four main types of reliability:

Type of reliability Measures the consistency of…
Test-retest The same test over time.
Interrater Different researchers or observers.
Parallel forms Two equivalent versions of the same test.
Internal consistency Items within a single test or questionnaire.
Tip
Make your research process easier with QuillBot AI Chat. Use it to clarify research concepts, evaluate your sampling methods, and think through questions about reliability, validity, and research ethics.

Test-retest reliability

Test-retest reliability measures whether the same test produces similar results when it’s given to the same participants at different points in time.

It’s most appropriate for measuring characteristics that are expected to remain stable, such as personality traits or cognitive ability.

The importance of test-retest reliability

Results can change because participants are distracted, tired, or influenced by temporary circumstances. Test-retest reliability helps determine whether a measurement remains consistent despite these factors.

How to measure test-retest reliability

Researchers administer the same test twice to the same participants and calculate the correlation between the two sets of scores. A stronger correlation indicates higher reliability.

Test-retest reliability example
A researcher creates a questionnaire for a descriptive study to measure people’s attitudes toward remote work, which are expected to remain relatively stable over a short period of time.

The same group of participants completes the questionnaire twice, with a few weeks between each test. However, the participants’ scores vary significantly between the two attempts, suggesting that the questionnaire has low test-retest reliability. The researchers decide to conduct semi-structured interviews to follow up on the results and gather some qualitative data.

How to improve test-retest reliability

  • Write clear, specific questions and tasks that are less likely to be affected by participants’ mood, concentration, or other temporary factors.
  • Keep your testing conditions as consistent as possible. Try to minimize external influences during data collection and ensure that every participant completes the test under similar conditions.
  • Choose the timing of repeated tests carefully. Keep in mind that participants may remember their earlier responses (recall bias) or genuinely change over time, and consider these factors when interpreting your results.

Interrater reliability

Interrater reliability measures how consistently different people evaluate or observe the same phenomenon.

It’s especially important in research that relies on human judgment, such as coding interviews, observing behavior, or rating performance.

The importance of interrater reliability

Different researchers may interpret the same situation differently. High interrater reliability shows that the results depend on the measurement process rather than individual opinions.

How to measure interrater reliability

Multiple researchers assess the same participants or observations. Their ratings are then compared using a statistical measure of agreement or correlation. Using multiple researchers is also called investigator triangulation.

Interrater reliability example
A company conducts a correlational study to evaluate the quality of customer service interactions. Several researchers review recorded customer calls and rate factors such as helpfulness, communication skills, and problem-solving ability using a predefined scoring system.

Because the researchers assign similar scores to the same calls, the rating system in this quantitative research demonstrates high interrater reliability.

How to improve interrater reliability

  • Clearly define your variables, illustrate the predicted relationships in a conceptual framework, and decide in advance how they will be measured. Make sure everyone involved understands exactly what is being observed, rated, counted, or categorized.
  • Create detailed and objective guidelines for evaluating the variables. Clear criteria help ensure that different researchers apply the same standards when making observations or assigning scores.
  • If multiple people are collecting or analyzing data, make sure the researchers receive the same instructions and training. This helps reduce differences in how they interpret the criteria and improves consistency between their results.

Parallel forms reliability

Parallel forms reliability evaluates whether two different versions of a test produce similar results.

The two versions should measure the same construct with questions that are equivalent in difficulty and content.

Importance of parallel forms reliability

If you plan to use multiple versions of the same test (for example, to prevent participants from simply remembering and repeating their previous answers), you first need to check that each version measures the same thing consistently and produces reliable results.

How to measure parallel forms reliability

The most common way to measure parallel forms reliability is to create a larger set of questions designed to measure the same concept and then divide those questions into two equivalent test versions.

The same group of participants completes both versions of the test, and the results are compared by calculating the correlation between the two sets of scores. A strong correlation between the results suggests that the two versions are measuring the same thing consistently, indicating high parallel forms reliability.

Parallel forms reliability example
A researcher develops two versions of a reading comprehension test to measure students’ language skills. The questions are divided into two similar sets, with both versions covering the same types of reading abilities and difficulty levels.

Through simple random sampling, 60 students are selected to complete both versions of the test, and their scores are compared. Because the results are very similar across both versions, the test demonstrates high parallel forms reliability.

How to improve parallel forms reliability

  • Make sure both versions of the test are designed to measure the same thing. Use the same underlying theoretical framework, learning objectives, or criteria when creating each set of questions.
  • Keep the two versions as similar as possible in terms of difficulty, format, and content. Differences between the tests should not affect participants’ scores.
  • Test both versions before using them in your research, especially when you’re conducting exploratory research. A pilot study can help identify loaded questions and questions that are unclear, too easy, too difficult, or inconsistent between the two versions.

Internal consistency

Internal consistency looks at how closely related the different items within a test are when they are designed to measure the same underlying concept. In other words, it checks whether the questions or statements in a questionnaire are all working together to assess the same construct.

Unlike some other types of reliability, internal consistency can be measured using data from a single test session. You don’t need to test participants again or involve multiple researchers, which makes it a useful option when you only have one set of data available.

Importance of internal consistency

When you create a questionnaire or rating scale where multiple items are combined into a single score, it’s important to check that each item is measuring the same underlying concept. If some questions produce results that conflict with the others or appear to measure something different, the overall test may not provide a reliable measure of the construct you’re trying to assess.

How to measure internal consistency

There are two common methods used to measure internal consistency:

  • Average inter-item correlation: This method looks at the relationship between all possible pairs of items designed to measure the same construct. Researchers calculate the correlation between each pair of items and then find the average correlation across all items. A higher average correlation suggests that the items are measuring the same underlying concept.
  • Split-half reliability: With this method, researchers divide a set of test items into two groups, usually at random. After participants complete the full test, the scores from both halves are compared. A strong correlation between the two sets of scores indicates that the test items are consistent with one another.
Internal consistency example
A researcher creates a questionnaire to measure employee job satisfaction. Through stratified sampling, 100 participants are selected to rate their agreement with statements such as “I feel valued at work” and “I enjoy my daily tasks” on a Likert scale.

If the questionnaire has high internal consistency, participants who score highly on one job satisfaction item should generally score highly on the other related items as well. However, if responses to the different questions show little connection with one another, this suggests that the questionnaire has low internal consistency.

How to improve internal consistency

  • Make sure all questions or items are clearly connected to the same concept. When designing a questionnaire or test, each item should be based on the same theory or research framework and should contribute to measuring the same underlying construct.
  • Review and refine items that do not fit with the rest of the test. Questions that measure a different concept or produce inconsistent responses can reduce the overall reliability of the measurement.
  • Use clear and specific wording when creating questions or rating scales. Well-designed items help participants interpret questions in the same way and provide more consistent responses.

Which type of reliability applies to my research?

Reliability is something to consider throughout your research process: from designing your research and collecting your data to analyzing your results and reporting your findings. The type of reliability you should assess depends on your research method, the type of data you collect, and how your measurements are being used.

Research situation Reliability type
Measuring a stable characteristic over time. Test-retest reliabilty
Multiple people collect or rate data. Interrater reliability
Comparing two versions of the same assessment. Parallel forms reliability
Using multiple questions to measure one construct. Internal consistency

Whenever it is relevant and possible, calculate a statistical measure of reliability and report it alongside your research results. Including reliability statistics helps readers understand how consistent your measurements are and gives them more confidence in the quality of your findings.

The statistical measure you use to assess reliability can depend on the type of data you are measuring, which is closely related to the level of measurement of your data and your research method.

  • For nominal data (categories with no order), reliability is often assessed using measures of agreement, such as Cohen’s kappa.
  • For ordinal data (ordered categories), reliability may be assessed using weighted kappa or ordinal versions of correlation measures.
  • For interval data or ratio data (continuous numerical measurements), reliability is often assessed using correlations, intraclass correlation coefficients, or measures like Cronbach’s alpha.
Tip
A reliable measurement produces consistent results, but you also need to consider whether it accurately measures what it is intended to measure. Different types of validity help you evaluate different aspects of your research. For example:

  • Construct validity examines whether your method accurately measures the concept you are studying.
  • Content validity looks at whether your measurement covers all relevant aspects of a concept
  • External validity considers whether your findings can be generalized to other contexts.

Considering both reliability and validity helps ensure that your research methods produce accurate, meaningful, and trustworthy results.

Frequently asked questions about reliability

What is a good inter-rater reliability score?

A good inter-rater reliability score depends on the statistic used and the context of the study.

For Cohen’s kappa (two raters), common guidelines are:

  • < 0.20: Poor agreement
  • 0.21–0.40: Fair agreement
  • 0.41–0.60: Moderate agreement
  • 0.61–0.80: Substantial agreement
  • 0.81–1.00: Almost perfect agreement

For the Intraclass Correlation Coefficient (interval or ratio data), similar thresholds are used:

  • < 0.50: Poor agreement
  • 0.51–0.75: Moderate agreement
  • 0.76–0.90: Good agreement
  • > 0.91: Excellent agreement
What is inter-rater reliability in psychology?

In psychology, inter-rater reliability refers to the degree of agreement between different observers or raters who evaluate the same behavior, test, or phenomenon. 

It ensures that measurements are consistent, objective, and not dependent on a single person’s judgment, which is especially important in research, clinical assessments, and behavioral studies.

High inter-rater reliability indicates that results are dependable and reproducible across different raters.

What is the formula for calculating inter-rater reliability?

There isn’t just one formula for calculating inter-rater reliability. The right one depends on your data type (e.g., nominal data, ordinal data) and the number of raters.

  • Cohen’s kappa (κ) is commonly used for two raters
  • Fleiss’ kappa is typically used for three or more raters
  • The Intraclass Correlation Coefficient (ICC) is used for continuous data (interval or ratio). This is based on analysis of variance (ANOVA)

The most commonly used formula (for Cohen’s kappa) is:
\kappa = \dfrac{{{P}_o}-{{P}_e}}{{1}-{P_e}}
Po is the observed proportion of agreement, and Pe stands for the expected agreement by chance.

Should I use a 5- or 7-point Likert scale?

Though traditional Likert scales include a 5-point response scale, some research has indicated that 7-point scales provide more reliable results.

As a rule of thumb, 5-point scales are better for unipolar constructs, which range from zero to positive, such as frequency. You may want to use 7-point scales for bipolar (or dichotomous) constructs that range from negative to positive, such as quality—some evidence suggests that doing so can increase reliability.

Other interesting articles

If you want to know more about colors, letters, or the meaning of emojis, make sure to check out some of our other articles with explanations and examples.
Is this article helpful?
Julia Merkus, MA

Julia has a bachelor in Dutch language and culture and two masters in Linguistics and Language and speech pathology. After a few years as an editor, researcher, and teacher, she now leads the Quillbot content team. She also writes articles about her specialist topics: grammar, linguistics, research, and statistics.

Join the conversation

Please click the checkbox on the left to verify that you are a not a bot.