Verbal operants are categories of language defined not by what words mean in a dictionary sense, but by why a person says them and what consequences maintain them. The concept comes from B.F. Skinner’s 1957 book Verbal Behavior, which proposed that every instance of language can be classified by the environmental conditions that evoke it and the type of reinforcement that follows. Rather than treating language as a single unified skill, Skinner broke it into functionally distinct units: a child saying “cookie” because she wants one is doing something fundamentally different from a child saying “cookie” because she sees one, even though the word is identical. That functional distinction is the core idea behind verbal operants, and it has shaped decades of language intervention, particularly for children with autism and other developmental disabilities.
The Primary Verbal Operants
Skinner identified several elementary verbal operants, each defined by a unique combination of what causes the behavior and what reinforces it. The two that have received the most research attention are the mand and the tact, though the others play equally important roles in fluent communication.
A mand is a request. It is evoked by a state of wanting or needing something, and it is reinforced by getting that specific thing. When you say “water” because you’re thirsty and someone hands you a glass, that’s a mand. The critical feature is that the speaker’s motivation drives the response, and the reinforcement directly matches that motivation. Research on mand training has emphasized the importance of ensuring that manding comes under the control of the learner’s genuine motivation rather than artificial prompts. If a child only says “water” when a therapist holds up a picture card, the word may be under the wrong kind of control and less likely to show up in everyday situations where the child is actually thirsty.
A tact is a label or comment about the environment. It is evoked by the presence of something the speaker can perceive and reinforced by social acknowledgment, not by receiving the thing itself. Saying “dog” when you see a dog walk by, or “it’s cold” when you step outside, are tacts. The reinforcement is typically generalized: someone nods, agrees, or engages in conversation. Empirical research has consistently supported Skinner’s claim that tacts and mands are functionally independent. A child who can label a cookie when she sees one does not automatically know how to ask for a cookie when she wants one. This is one of the framework’s most practically important findings.
An echoic is a vocal imitation. When someone says “ball” and you repeat “ball,” that’s an echoic. The stimulus is someone else’s speech, and the response matches it in sound. Echoics are foundational for early language learning because they allow a listener to practice new words. Researchers have proposed that echoic behavior, along with other responses that share formal similarity with their stimulus (like imitating a sign in sign language), belongs to a broader category called “duplic” behavior, where the response looks or sounds like the model.
An intraverbal is a verbal response to someone else’s verbal behavior, but without the point-to-point match you see in an echoic. If someone says “How are you?” and you say “Fine,” that’s an intraverbal. So is answering “What’s the capital of France?” with “Paris.” There’s no physical resemblance between the question and the answer. Intraverbals are the backbone of conversation, academic responding, and social interaction. Analysis of intraverbal behavior has identified multiple levels of complexity in how verbal stimuli control responses, from simple one-word associations to conditional discriminations where the correct answer depends on context.
Textual behavior is reading aloud: seeing written text and producing the corresponding spoken words. It involves a point-to-point correspondence between the written stimulus and the vocal response, but no formal similarity, since printed letters don’t look like the sounds they represent. This places textual behavior in what some researchers call the “codic” category, distinguishing it from echoic (duplic) behavior. The reverse process, hearing speech and writing the corresponding words, is called transcription or dictation-taking.
Why the Same Word Can Be Different Operants
The feature that makes verbal operants confusing at first, and powerful once you understand them, is that the same word can function as completely different operants depending on the circumstances. A child who says “ball” might be manding (she wants the ball), tacting (she sees a ball), echoing (her teacher just said “ball”), or intraverbing (someone asked “What do you play catch with?”). Four instances of the same word, four different functional categories, each potentially learned independently.
This matters enormously in practice. Research reviews have found that language training programs may fail when they don’t provide explicit training across each elementary verbal operant and across both speaker and listener repertoires separately. Teaching a child to label objects (tact) does not guarantee the child can request those objects (mand), answer questions about them (intraverbal), or follow instructions involving them (listener behavior). Each functional use often needs to be taught on its own terms.
Autoclitics and Multiply Controlled Behavior
Beyond the elementary operants, Skinner described autoclitics, which are verbal responses that modify or qualify other verbal behavior. When you say “I think it might rain,” the words “I think” and “might” are autoclitics. They don’t label rain or request anything; they signal uncertainty about the rest of the sentence. Autoclitics are how speakers build grammar, qualify assertions, negate statements, and organize their speech into sentences rather than single words. They act on the listener by adjusting how the primary verbal operant is received.
Real-world speech is rarely a “pure” example of any single operant. Most utterances involve multiple controlling variables at once. Saying “Can I have that red ball?” is partly a mand (you want the ball), partly a tact (you’re describing its color), and partly autoclitic (the sentence structure guides the listener’s response). Skinner called these multiply controlled verbal operants, and researchers have applied this concept to the design of communication training programs, including the Picture Exchange Communication System (PECS), which is widely used with children who have limited vocal speech. The PECS training sequence is designed to progressively establish multiply controlled verbal behavior, moving learners from simple mands to more complex utterances that combine requesting and describing.
Transfer Across Operants
One of the most clinically relevant questions in verbal operant research is whether learning a word in one operant category helps a person use it in another. The short answer: sometimes, but not reliably without explicit teaching.
Transfer procedures are structured teaching methods designed to move stimulus control from one operant to another. A common approach starts with an echoic (the teacher says a word, the child repeats it), then shifts to a tact (the child sees the object and says the word without a vocal model). In one applied study, a combination of receptive-to-echoic-to-tact transfer procedures was used during brief instructional sessions with a child with autism. Without the teaching procedure, the child acquired no new tacts. With the procedure in place, the child acquired thirty new tacts over sixty sessions.
Transfer from tact to mand has also been investigated. In research with adults with severe intellectual disabilities, mands for certain items emerged after those items had been taught as tacts, but only when the participants already had a minimal mand repertoire in place. In other words, having some experience with requesting in general seemed to facilitate the transfer. The takeaway for practitioners is that transfer is possible but not automatic, and building a basic repertoire in one operant can create a foundation for learning in another.
Mand Training and Challenging Behavior
Much of the applied research on verbal operants has focused on mands, for good reason. When people lack effective ways to request what they need, they often turn to challenging behavior instead: hitting, screaming, self-injury, or property destruction. Functional communication training (FCT) addresses this by teaching an alternative communicative response that serves the same function as the problem behavior.
There is a subtle but important distinction between FCT and mand training more broadly. FCT typically focuses on teaching a single response to quickly replace problem behavior, while mand training more often targets multiple responses to expand a person’s overall communication repertoire. Both draw on the verbal operant framework, but their goals differ. FCT prioritizes behavior reduction; mand training prioritizes communication growth.
Research has explored how to make mand training more effective within FCT. One study compared different reinforcement schedules during FCT for individuals with autism and found that requiring variety in mand responses (using what’s called a lag schedule, where each response has to differ from the one before) increased the diversity of communication while keeping challenging behavior low. This is a meaningful finding because a common criticism of FCT is that it can produce robotic, repetitive communication. Requiring varied responses pushes the learner toward more natural-sounding language.
Researchers have also examined whether it’s better to train new mands or use mands the person already knows during FCT. Both approaches have been shown to decrease problem behavior, but the choice can depend on the individual’s existing repertoire and how quickly problem behavior needs to be reduced.
Motivating Operations and the Problem of Artificial Control
A recurring concern in mand training is making sure the behavior is genuinely motivated rather than artificially prompted. If a therapist always holds up a juice box before asking “What do you want?”, the child may learn to say “juice” whenever the box appears, even if she isn’t thirsty. That response would look like a mand but technically be under discriminative control (triggered by the visual cue) rather than motivational control (triggered by wanting juice).
This distinction has practical consequences. Research on teaching mands for information in preschoolers with autism found that when alternative reinforcers like tokens were used instead of the naturally motivating consequence, the resulting behavior was less likely to generalize to new situations. A child who learns to ask “Where is it?” only because she gets a sticker for asking, rather than because she genuinely needs to find something, may not ask that question at home when she’s actually looking for her shoes.
Procedures that manipulate motivating operations, like creating situations where the child genuinely needs information or an item, produce more robust and generalizable manding. One study used rolling time-delay and prompt-fading procedures to successfully transfer mand control to the relevant motivating operation in children with autism, freeing the mands from the artificial controls that are common in clinical settings.
Assessing Verbal Operants in Practice
Clinicians working with children with autism and other developmental delays need structured ways to measure verbal operant repertoires. The most widely used tool for this purpose is the Verbal Behavior Milestones Assessment and Placement Program (VB-MAPP), which maps a child’s skills across manding, tacting, echoic behavior, intraverbal behavior, listener responding, and other domains. The milestones section of the VB-MAPP identifies where a child currently functions and what instructional goals should come next.
The VB-MAPP has been examined for inter-rater agreement, meaning whether two different assessors come to similar conclusions when evaluating the same child. Research has found that agreement varies across different sections of the assessment, which matters for program planning. If two therapists evaluate the same child and disagree substantially on intraverbal skills, for instance, they might write very different goals. Clinicians using the VB-MAPP are generally trained in its administration to minimize this variability.
The Chomsky Critique and Its Aftermath
No discussion of verbal operants is complete without acknowledging the most famous attack on the framework. In 1959, linguist Noam Chomsky published a review of Skinner’s Verbal Behavior that has been called one of the most influential documents in the history of psychology. Chomsky argued that Skinner’s functional categories couldn’t account for the creativity and complexity of human language, particularly children’s ability to produce sentences they’ve never heard before.
The review had an outsized impact. For decades, it effectively sidelined Skinner’s analysis in mainstream linguistics and cognitive psychology. Multiple rejoinders have been published over the years, but their influence has been limited compared to the original critique. Interestingly, researchers who have examined the history of Skinner’s analytical approach have noted that his functional classification of verbal operants corresponds reasonably well with formal analyses of languages, suggesting the two traditions are less incompatible than the debate implied.
In applied settings, the debate matters less than it does in theoretical circles. Clinicians working with children who have limited language aren’t trying to explain universal grammar; they’re trying to teach functional communication. For that purpose, the verbal operant framework has proven remarkably useful. A growing body of empirical research supports many of the tenets of Skinner’s conceptualization, even as many areas remain underexplored.
Relational Frame Theory as an Extension
One of the more interesting modern developments is the attempt to synthesize Skinner’s verbal operants with Relational Frame Theory (RFT), a newer behavioral account of language that emphasizes derived relational responding. RFT researchers have examined each of Skinner’s verbal operants and identified two types within each: one based on direct contingencies of reinforcement (the kind Skinner described) and another based on arbitrarily applicable relational responding, which is the ability to relate stimuli in flexible ways without direct training.
For example, if a child learns that a spoken word relates to an object (tacting), and then learns that a written word relates to the spoken word (textual behavior), the child may derive a new relation between the written word and the object without being explicitly taught. This kind of derived responding goes beyond what Skinner’s original framework easily explains, and RFT provides a behavioral account of how it happens. The synthesis is still being worked out, and not all behavior analysts accept it, but it represents the most serious attempt to extend the verbal operant framework to cover the generative aspects of language that Chomsky argued Skinner missed.
Private Events and Covert Verbal Behavior
Skinner treated thinking as covert verbal behavior, verbal operants that occur at such a reduced scale that they aren’t observable to anyone except the person doing the thinking. This is one of the more philosophically provocative claims in the framework. When you silently rehearse what you’re going to say at a meeting, you’re engaging in covert echoic and intraverbal behavior. When you mentally label something you see, that’s a covert tact.
The challenge, as behavioral researchers have discussed, is that private stimuli can only acquire the ability to control verbal responses when they correlate with public stimuli that the verbal community can observe and reinforce. You learn to say “I’m anxious” because there are publicly observable signs of anxiety (sweating, fidgeting) that others can see and use to teach you the word. The accuracy of your private verbal behavior is therefore limited by how well your internal states map onto external indicators that others can respond to. This means self-knowledge, from a verbal operant perspective, is always somewhat imprecise, shaped by the crudeness of the teaching conditions rather than by some inherent limitation of introspection.
Communication Systems Beyond Vocal Speech
Verbal operants are not limited to spoken words. Skinner defined verbal behavior by its social function, not its physical form. Signing, pointing to picture symbols, typing on a speech-generating device, or exchanging picture cards all count as verbal behavior if they are reinforced through the mediation of another person. This definitional flexibility has been important for augmentative and alternative communication (AAC) systems.
The Picture Exchange Communication System (PECS) is one example where the verbal operant framework has been directly applied to a non-vocal communication modality. PECS begins by teaching simple mands: a child hands a picture of a desired item to a communication partner and receives the item. Later phases introduce discriminating between pictures, constructing sentence strips, and responding to questions, progressively building tact and intraverbal repertoires. Researchers have described how PECS’s training sequence is designed to establish multiply controlled verbal behavior, moving from single operants to utterances where multiple variables are at work simultaneously.
The framework also raises a useful caution about AAC. Some researchers have proposed refining the category labels for non-vocal modalities, since terms like “echoic” and “textual” imply specific sensory channels. Imitating a sign doesn’t involve the auditory-vocal channel that “echoic” implies, and reading Braille doesn’t involve the visual channel that “textual” implies. Proposed broader categories, “duplic” for any response that formally matches its model and “codic” for any response with point-to-point correspondence but no formal similarity, handle these cases more cleanly.
Where the Evidence Is Thin
Research on verbal operants is unevenly distributed. Reviews of the empirical literature have consistently found that most studies focus on mands and tacts, while echoic, intraverbal, autoclitic, and textual behavior receive less attention. This means the claims about the less-studied operants rest more on Skinner’s theoretical analysis and less on controlled experimental evidence. The research that does exist supports Skinner’s core distinction between the operant categories, but the field has acknowledged that many areas of verbal behavior research have yet to be adequately addressed.
Neuroimaging research has begun to explore how operant learning maps onto brain activation. Studies using functional neuroimaging have found activation in frontal and striatal brain regions during the presentation of discriminative stimuli, which is consistent with what would be expected from a reinforcement-learning perspective. But this line of research is still in early stages and has not yet produced findings specific to verbal operants as distinct from other operant behavior. The gap between the clinical utility of the verbal operant framework and the neuroscience behind it remains wide.

