An ABA reversal design is a single-case experiment that tests whether an intervention actually causes a change in behavior by introducing it, removing it, and watching what happens. The name comes from the sequence of phases: A stands for the baseline (no intervention), B stands for the treatment, and the return to A is the “reversal” that gives the design its power. If behavior improves during B, worsens when B is removed, and improves again when B is reintroduced in an extended ABAB version, the case for a cause-and-effect relationship becomes hard to dismiss. The design is a workhorse of applied behavior analysis and special education research, but it carries assumptions and limitations that matter in practice.
How the Phases Work
The simplest version uses three phases. In the first A phase, a researcher observes and records behavior under normal conditions, with no intervention in place. This baseline establishes what behavior looks like without any manipulation. In the B phase, the intervention is introduced and behavior is tracked using the same measurement. In the second A phase, the intervention is pulled away and baseline conditions return. If the behavior shifts in the expected direction during B and then shifts back toward baseline levels during the second A phase, the researcher has initial evidence that the intervention, not some coincidence, drove the change.
Most published studies extend the sequence to ABAB, adding a second intervention phase. That extra B phase does two things. It provides a built-in replication of the treatment effect within the same participant, and it ends the study with the intervention in place, which is often the ethical thing to do when the intervention helps. A researcher studying a classroom strategy for reducing disruptive behavior, for instance, would not want to end the study with the student back in baseline conditions if the strategy was working.
Why Removing the Intervention Is the Key Move
The return-to-baseline phase is what separates this design from a simple before-and-after comparison. Without it, any change observed during the B phase could be explained by maturation, practice effects, seasonal changes, or just the passage of time. The reversal phase controls for those possibilities: if the behavior worsens specifically when the intervention is removed and then improves again when it comes back, the most straightforward explanation is that the intervention itself is responsible.
A large-scale analysis of 501 ABAB datasets drawn from theses and dissertations found that initial treatment effects observed in the first AB pair were successfully replicated in subsequent phase reversals in roughly 85% of cases. The same analysis found that larger effect sizes in the initial AB comparison predicted a higher probability of successful replication, which makes intuitive sense: a strong, clear effect is more likely to show up again than a subtle one.
1PubMed Central. Using Single-Case Designs in Practical Settings: Is Within-Subject Replication Always Necessary?That 85% figure is reassuring, but it also means that in about 15% of cases, the initial effect did not cleanly replicate. Behavior is messy. A child who learns a new skill during the B phase may retain some of it during the withdrawal, or environmental changes unrelated to the study may shift baseline conditions. The reversal is the design’s strength, but it does not guarantee a tidy demonstration every time.
Withdrawal Versus Reversal, and Why the Terms Get Confused
The terms “withdrawal design” and “reversal design” originally referred to two different things. In a withdrawal design, the researcher simply stops delivering the intervention during the return-to-baseline phase. In a true reversal design, the researcher actively applies the intervention to a different behavior or switches the reinforcement contingencies so that the original target behavior is no longer reinforced. The distinction matters conceptually because the reversal version provides an even stronger demonstration of experimental control: not only does behavior change when the intervention is removed, it changes in a specific, predicted direction when the contingencies are flipped.
In practice, though, the field has largely stopped making this distinction. A review of major behavior-analytic journals published between 2009 and 2013 found a strong preference among researchers for the term “reversal” regardless of whether the study actually reversed contingencies or simply withdrew the intervention.
2Behavioral Interventions. Withdrawal Versus Reversal: A Necessary Distinction?If you are reading published studies, expect to see “ABAB reversal design” used as a blanket label for any design that follows the A-B-A-B phase structure, whether or not a true reversal of contingencies occurred. Knowing the original distinction can help you evaluate a study more carefully, because a genuine reversal of contingencies is a more rigorous demonstration than a simple withdrawal, but you should not assume the label tells you which one was actually done.
How Researchers Read the Data
The traditional method for analyzing ABAB data is visual analysis: a trained researcher looks at a graph of the data points across phases and judges whether the pattern shows a clear, consistent change that coincides with the introduction and removal of the intervention. This sounds subjective, and it can be. Different analysts looking at the same graph sometimes reach different conclusions, particularly when effects are moderate rather than dramatic.
To address this, researchers have developed structured protocols that walk the analyst through a systematic set of judgments about level, trend, and variability within and across phases, ultimately producing a numeric score that summarizes the strength of the evidence.
3PubMed Central. Systematic Protocols for the Visual Analysis of Single-Case Research DataThese protocols do not replace visual analysis so much as impose consistency on it. The analyst still looks at the graph, but the structured questions reduce the chance of two analysts interpreting the same data differently.
Alongside visual analysis, quantitative effect-size indices have gained traction. One widely used measure is Tau-U, which combines two pieces of information: the degree of non-overlap between baseline and treatment data points and any trend within phases.
4PubMed. Combining nonoverlap and trend for single-case research: Tau-UOther common indices include PEM (Percentage of Data Points Exceeding the Median) and PAND (Percentage of All Non-Overlapping Data). Each captures a slightly different aspect of the effect, and researchers sometimes calculate all three to see whether they converge on the same conclusion.
In one study applying these indices to a single-case design, Tau-U values ranged from 0.91 to 0.95 across participants, PEM hit the maximum of 100% for all participants, and PAND values ranged from 94% to 97%, all indicating very large effects.
5Research Square. Examining the Role of Scientific Inquiry in Environmental Science Education: A Single-Case Study on Conceptual ChangeNumbers that high are not typical of every ABAB study; that example reflects a case where the intervention produced a large, clear change. But it illustrates how quantitative indices complement the visual impression. When visual analysis suggests a strong effect and three different quantitative measures agree, the conclusion is harder to second-guess.
The choice among effect-size indices is not trivial. Different indices can produce different classifications of the same data. A comparison of PND, IRD, and Tau-U across 111 baseline-to-treatment comparisons drawn from positive behavior support studies found disagreements in how each index classified effect levels, highlighting that the choice of index can influence conclusions.
6Korean Association for Behavior Analysis. A Comparison of Effect Size Indices for Single-Case Research on Positive Behavior Support in Inclusive Education Settings: PND, IRD, and Tau-UWhere the Design Shows Up in Practice
ABAB reversal designs are most commonly associated with applied behavior analysis, particularly in work with individuals with autism spectrum disorder and other developmental disabilities. The design is a natural fit for this population because interventions are often individualized, and the question is not whether a treatment works on average across a large group but whether it works for this particular person.
One study used an ABAB reversal design to test video self-modeling delivered on an iPad for a secondary student with autism and intellectual disability during science instruction. The student’s correct academic responses increased during the intervention phases and decreased when the intervention was removed, providing a clear demonstration of the treatment effect within a single individual.
7Education and Training in Autism and Developmental Disabilities. Using Video Self-Modeling Via iPads to Increase Academic Responding of an Adolescent with Autism Spectrum Disorder and Intellectual DisabilityA separate study used the same ABAB structure to compare instructional delivery via iPad versus traditional materials for two students with autism. Both students showed lower levels of challenging behavior and higher academic engagement during the iPad phases, with the pattern reversing during the traditional-materials phases.
8Research in Autism Spectrum Disorders. The effect of instructional use of an iPad® on challenging behavior and academic engagement for two students with autismThese studies illustrate how the ABAB design captures the individual-level causal story that group-average designs would obscure. A randomized controlled trial might tell you an intervention works for most students; an ABAB study tells you it worked for this student, right here, during these sessions.
The design is also used outside of disability research. Any behavioral or educational intervention that can be cleanly introduced and removed is a candidate. Researchers have used reversal designs to evaluate classroom management strategies, organizational behavior interventions, exercise programs, and environmental modifications. A study on overcorrection procedures, for example, used a reversal design to compare different durations of a behavioral consequence, finding that longer durations did not produce the superior effects that had been previously claimed.
9Behavior Modification. Parametric Analysis of Overcorrection Duration EffectsWhen the Design Does Not Work
The ABAB design has a fundamental requirement: the behavior must be reversible. If an intervention teaches a skill that the person retains even after the intervention is removed, the second A phase will not show a return to baseline. That does not mean the intervention failed. It means the design is a poor fit for that type of intervention. You would not use an ABAB design to study whether teaching a child to read produces lasting reading ability, because reading does not disappear when the instruction stops. The lack of reversal would be a success for the learner but a failure for the experimental logic.
This is one of the main reasons researchers sometimes prefer alternative single-case designs. A multiple-baseline design, for instance, introduces the intervention at staggered times across different participants, settings, or behaviors. If the change happens in each case only after the intervention is introduced, the causal case is strong, and no withdrawal is needed. For skills that are expected to stick, or for situations where removing an effective intervention would be ethically problematic, multiple-baseline designs are usually the better choice.
Another limitation involves carryover effects. If the B phase changes something about the environment or the person that persists into the second A phase, the reversal will be incomplete or absent even though the intervention truly caused the change. A medication that alters brain chemistry over weeks, for example, would not wash out cleanly during a brief return to baseline. The ABAB design works best when the intervention’s effects are expected to be relatively immediate and relatively temporary when removed.
Practical constraints matter too. In applied settings like classrooms or clinics, maintaining clean phase transitions is not always straightforward. Teachers may drift in implementation fidelity. Parents may continue using intervention strategies at home even during a supposed withdrawal. The real world does not always cooperate with the sharp phase boundaries the design requires.
How Quality Standards Evaluate These Studies
Single-case designs, including the ABAB, have historically been viewed with some skepticism by evidence-clearinghouse bodies accustomed to large randomized trials. The What Works Clearinghouse, which reviews educational interventions in the United States, has developed specific standards for evaluating single-case research. These standards address minimum phase lengths, the number of data points per phase, and whether the design demonstrates at least three instances of an effect at different points in time (which an ABAB design satisfies by producing effects at the A-to-B transitions and losses at the B-to-A transitions).
These standards have been critiqued and updated over time. One review recommended that visual analysis be used to verify a study’s design standards and operational issues, identified limitations in certain statistical measures used to compare single-case effect sizes to group-design effect sizes, and called for the inclusion of diversity and equity considerations in future standards.
10PubMed. Single-case design standards: An update and proposed upgradesThe field is still working out how to translate single-case evidence into the kind of standardized scores that policymakers and funding agencies want to see, but the direction of travel is toward treating well-conducted ABAB and other single-case designs as legitimate evidence rather than mere case reports.
Validity threats familiar from group research apply to single-case studies too, though they play out differently. A framework for mapping standard validity-threat categories onto single-case experiments has been developed specifically to help behavior analysts communicate the rigor of their designs to audiences outside the field.
11Springer. Applying the Taxonomy of Validity Threats from Mainstream Research Design to Single-Case Experiments in Applied Behavior AnalysisThreats like history (something outside the study changes between phases), maturation (the participant develops over time), and testing effects (repeated measurement alters behavior) all apply. The ABAB structure mitigates some of these threats better than a simple AB comparison does, because the repeated phase changes make it unlikely that an outside event would coincidentally align with every transition. But the design does not eliminate all threats, and careful researchers document potential confounds rather than assuming the phase structure handles everything.
How ABAB Compares to Other Single-Case Designs
The ABAB reversal design is one member of a family of single-case experimental designs, each suited to different questions. The alternating treatments design rapidly switches between two or more conditions within the same participant, sometimes within the same session, to compare their relative effects.
12PubMed Central. Alternating treatments design: one strategy for comparing the effects of two treatments in a single subjectWhere an ABAB design asks “does this intervention work?”, an alternating treatments design asks “which of these interventions works better?” The two designs answer fundamentally different questions, so the choice between them depends on what the researcher or clinician needs to know.
Multiple-baseline designs, as mentioned earlier, stagger the introduction of an intervention across participants, behaviors, or settings. They avoid the need for withdrawal entirely, which makes them the default when reversal is impractical or unethical. Their weakness is that they require multiple baselines to be independent of one another. If introducing an intervention for one participant somehow affects another participant’s baseline, the logic breaks down.
Changing-criterion designs set progressively higher performance targets across phases, demonstrating experimental control by showing that behavior matches each new criterion. These are useful for shaping behaviors like exercise duration or academic output, where the goal is gradual improvement rather than the presence or absence of an effect.
Each design has trade-offs. The ABAB is often the most convincing when it works, because the repeated introduction and removal of the intervention within the same person is a powerful demonstration. But “when it works” is doing a lot of heavy lifting in that sentence. Many real-world interventions are not cleanly reversible, and many applied settings do not permit the kind of phase control the design demands. Experienced researchers choose the design that fits the question and the context rather than defaulting to one structure.
Ethical Tensions in Removing Effective Treatments
The most persistent criticism of the ABAB design is an ethical one: if the intervention is helping, why would you take it away? In clinical practice, deliberately returning a client to a less effective condition to demonstrate experimental control feels uncomfortable at best and harmful at worst. A child whose aggressive behavior drops dramatically during an intervention phase is going to have a harder time, and so will the people around them, during a withdrawal phase.
Researchers navigate this tension in several ways. One is to keep the withdrawal phase as short as possible, collecting just enough data to establish that behavior has begun to shift back toward baseline levels before reintroducing the intervention. Another is to use the ABAB design only for interventions where the withdrawal is unlikely to cause serious harm, reserving more powerful but ethically fraught applications for multiple-baseline designs that avoid withdrawal altogether.
There is also a practical argument in favor of the withdrawal phase. If a practitioner is going to recommend an intervention for long-term use, knowing that the intervention is actually responsible for the improvement, rather than coincidence, matters. A brief withdrawal that confirms the causal link may justify sustained investment of time, money, and effort that would otherwise rest on shaky evidence. The discomfort of a temporary withdrawal can be weighed against the cost of committing to an intervention that was never really working.
Some researchers have pushed back on the assumption that within-subject replication is always necessary. The analysis of 501 ABAB datasets mentioned earlier found that in the vast majority of cases, the initial AB result held up, suggesting that for large, clear effects, a practitioner in a clinical setting might reasonably trust the initial AB finding without insisting on a full ABAB sequence.
13PubMed Central. Using Single-Case Designs in Practical Settings: Is Within-Subject Replication Always Necessary?That is a pragmatic compromise, not a methodological one. The full ABAB still provides stronger evidence. But in applied settings where the priority is helping a client rather than publishing a paper, the calculus may favor moving forward with a promising result rather than engineering a withdrawal for the sake of replication.

