What IOA actually measures

Agreement is not accuracy. Knowing the difference is what stops a high score being over-read.

Interobserver agreement (IOA) is the degree to which two independent observers, watching the same behaviour at the same time, produce the same record. It is a check on whether your measurement system is repeatable — nothing more.

That distinction matters because a high score is easy to misread. Two observers can agree perfectly and both be wrong, if they share the same misunderstanding of the operational definition. Agreement tells you the definition is being applied consistently; it does not tell you it is being applied correctly, and it says nothing at all about whether the intervention is working.

When agreement comes back low, the fault is almost never in the observers' attention. It is in the definition. If two trained people watching the same child disagree about whether something counted, the definition has left room for judgement — and the fix is to tighten it, not to retrain the observers on a vague target.

Choosing the right method

The method is decided by how the data were collected, not by which number you would prefer to report.

If you collected… Use
A single tally for the whole session Total count IOA
Counts within each interval Exact count-per-interval or mean count-per-interval
Occurrence / non-occurrence per interval Interval-by-interval
Interval data, low-rate behaviour Scored-interval IOA
Interval data, high-rate behaviour Unscored-interval IOA
Discrete trials with a defined response per trial Trial-by-trial IOA
How long the behaviour lasted Total duration or mean duration-per-occurrence

Every method below reduces to the same idea — agreement divided by opportunity, multiplied by 100. What changes is the definition of an agreement and what counts as an opportunity.

Count-based methods

For frequency and event recording — the tally-mark family.

Total count IOA

smaller count ÷ larger count × 100

The simplest and the weakest. It compares two session totals and ignores everything about when the behaviour occurred.

Worked example. Observer 1 records 12 instances of calling out; Observer 2 records 15. 12 ÷ 15 = 0.80, so IOA is 80%.

The weakness is real: two observers can produce identical totals while recording entirely different events. Use it when a session tally is genuinely all you have, and prefer an interval method when you have the choice.

Exact count-per-interval IOA

intervals with identical counts ÷ total intervals × 100

The most conservative count-based method. An interval only counts if both observers recorded exactly the same number.

Worked example. Across four intervals, Observer 1 records 2, 3, 0, 5 and Observer 2 records 2, 4, 0, 3. The counts match exactly in interval 1 (2 and 2) and interval 3 (0 and 0) — two of four intervals. IOA is 50%.

Mean count-per-interval IOA

sum of each interval's agreement ÷ number of intervals × 100

Calculate the smaller-over-larger ratio for every interval, then average them. An interval where both observers recorded zero is scored as full agreement.

Worked example. Using the same data as above — 2/3/0/5 against 2/4/0/3:

  • Interval 1: 2 ÷ 2 = 1.00
  • Interval 2: 3 ÷ 4 = 0.75
  • Interval 3: both zero = 1.00
  • Interval 4: 3 ÷ 5 = 0.60

The sum is 3.35; divided by 4 intervals that is 0.8375, so IOA is 83.8%.

Note what just happened: identical data returned 50% by the exact method and 83.8% by the mean method. Neither is wrong — but reporting a score without naming the method makes it uninterpretable.

Interval methods

For interval recording and time sampling — including data from the interval recording timer.

All three interval methods below use the same worked dataset, ten intervals scored for occurrence (1) and non-occurrence (0):

Interval 12345 678910
Observer 1 10110 01101
Observer 2 10100 01111

The observers disagree in exactly two places: interval 4 and interval 9.

Interval-by-interval IOA

agreements ÷ total intervals × 100

Also called point-by-point or total-interval IOA. Every interval counts, and an agreement is any interval where both observers recorded the same thing — including both recording nothing.

Worked example. Eight of the ten intervals match. 8 ÷ 10 = 80%.

This is the default for interval data, and it has a known inflation problem: for a behaviour that rarely occurs, most intervals are blank for both observers, and all that shared blankness counts as agreement.

Scored-interval IOA

agreements ÷ intervals either observer scored an occurrence × 100

The correction for low-rate behaviour. Only intervals where at least one observer recorded an occurrence are considered; intervals both left blank are discarded.

Worked example. At least one observer scored an occurrence in intervals 1, 3, 4, 7, 8, 9 and 10 — seven intervals. The observers agree in five of them (1, 3, 7, 8, 10). 5 ÷ 7 = 0.714, so IOA is 71.4%.

Unscored-interval IOA

agreements ÷ intervals either observer scored a non-occurrence × 100

The mirror image, used for high-rate behaviour where near-continuous occurrence would otherwise inflate the score.

Worked example. At least one observer scored a non-occurrence in intervals 2, 4, 5, 6 and 9 — five intervals. They agree in three (2, 5, 6). 3 ÷ 5 = 60%.

One dataset, three defensible answers: 80%, 71.4%, and 60%. This is why the method belongs next to the number every time you report it.

Trial-by-trial IOA

trials agreed ÷ total trials × 100

For discrete trial data, where each trial has one scored response. Structurally the same as interval-by-interval, with trials in place of intervals.

Worked example. Across 10 trials, two observers score the response the same way on 9. 9 ÷ 10 = 90%.

Duration methods

For behaviour measured in time rather than counts.

Total duration IOA

shorter total duration ÷ longer total duration × 100

The duration equivalent of total count, and it carries the same weakness — matching totals can conceal completely different timings.

Worked example. Observer 1 records 340 seconds of off-task behaviour; Observer 2 records 400. 340 ÷ 400 = 85%.

Mean duration-per-occurrence IOA

sum of each occurrence's agreement ÷ number of occurrences × 100

Compare the two timings for each individual occurrence, then average. More demanding than total duration and much harder to pass by coincidence.

Worked example. Three occurrences:

  • Occurrence 1: 30s and 35s → 30 ÷ 35 = 0.857
  • Occurrence 2: 60s and 60s → 1.000
  • Occurrence 3: 20s and 28s → 20 ÷ 28 = 0.714

The sum is 2.571; divided by 3 occurrences that is 0.857, so IOA is 85.7%.

Interpreting the score

What the benchmarks mean, and the exam question that trips people up.

80% is the conventional floor for treatment data. 90% or above is what research reporting and clinically risky behaviours call for. Below 80%, the honest reading is that you do not yet have data worth analysing — and the repair work belongs in the definition and in observer training, not in the intervention.

Collect agreement across at least 20% to 33% of sessions, spread through every phase rather than bunched at the beginning. Agreement measured only during baseline tells you nothing about whether observers drifted apart once the intervention started.

“What does a high IOA score indicate?”

This one appears on RBT and BCBA exams with four options that all sound plausible. The answer is that observers are consistently recording the behaviour the same way.

The distractors are each a different confusion:

  • The observer is using reinforcement correctly — that is treatment integrity, a separate measure of whether the procedure was delivered as written.
  • The technician implemented the plan correctly — treatment integrity again. IOA scores the measurement, not the intervention.
  • The procedure is no longer needed — that is a decision drawn from the behaviour data themselves, and it depends on reliable measurement rather than being indicated by it.

The underlying point is the one from the top of this page: agreement is a property of your measurement system, not of the learner, the plan, or the outcome.

Frequently asked questions

What does a high IOA score indicate?

That two independent observers recorded the same behaviour the same way. It tells you the operational definition is clear and both observers are applying it consistently. It does not tell you the data are accurate, that the intervention worked, or that the plan was implemented correctly — those are separate questions.

What is the formula for IOA?

There is no single formula. Every IOA method divides agreement by opportunity and multiplies by 100, but what counts as an agreement depends on how the data were collected. Total count IOA divides the smaller count by the larger; interval methods divide agreements by the number of intervals; duration methods divide the shorter duration by the longer.

What is a good IOA score?

80% is the conventional minimum for treatment data, and 90% or above is expected for research and for behaviours that carry clinical risk. Below 80%, treat the data as unreliable and revisit the operational definition or observer training rather than the intervention.

Which IOA method should I use?

Match the method to how you collected the data. Frequency tallies use total count or count-per-interval; interval recording and time sampling use interval-by-interval, scored-interval, or unscored-interval; discrete trial data use trial-by-trial; duration data use total duration or mean duration-per-occurrence.

Why did two IOA methods give me different scores from the same data?

Because they count different things. Interval-by-interval IOA includes intervals where both observers recorded nothing, which inflates agreement for low-rate behaviour. Scored-interval IOA excludes those, so it returns a lower and more conservative figure. Reporting the method alongside the score is what makes it interpretable.

How often should IOA be collected?

The common convention is a minimum of 20% to 33% of sessions, spread across conditions and phases rather than clustered at the start. Collecting it only during baseline tells you nothing about whether observer drift crept in once the intervention began.

Related tools & terms

Everything referenced above, in one place.

IOA Calculator
Runs total count, interval-by-interval, and partial interval agreement in the browser, with an interval-by-interval breakdown showing where observers diverged.
Interval Recording Timer
Momentary, partial-interval, and whole-interval recording with CSV export — the data these interval formulas expect.
Operational Definition
Where low agreement almost always originates, and the first thing to repair when a score comes back under 80%.
Frequency Tally Template
A printable grid for the count data that total count and count-per-interval agreement are calculated from.
Browse all resources →