What Is the Definition of Hate Speech?

Hate speech is expression that attacks, demeans, or incites hostility against people based on characteristics such as race, religion, sex, sexual orientation, disability, or age. That much is broadly agreed upon, but the details of where to draw the line vary enormously depending on who is drawing it. No single legal or academic definition has achieved universal acceptance, and the gaps between definitions are not trivial: they determine what gets removed from a social media platform, what triggers a criminal prosecution, and what remains protected expression. Understanding why the definition keeps shifting is just as important as knowing what the term generally means.

What Most Definitions Share

Despite disagreements at the edges, virtually every serious attempt to define hate speech includes a few core elements. The speech must target individuals or groups on the basis of a protected characteristic. It must go beyond simple rudeness or disagreement. And it typically involves threats, dehumanization, or denigration of the targeted group, along with either the intent to stir up hatred or the realistic potential to cause harm.1Journal of Language Aggression and Conflict. Understanding and appraising ‘hate speech’ Those criteria sound straightforward, but each one introduces ambiguity in practice. What counts as a “protected characteristic” differs by jurisdiction. Whether a statement qualifies as denigration or merely strong opinion is frequently contested. And whether intent matters, or only the likely effect on the audience, is a question that legal scholars, platform designers, and civil-rights advocates answer differently.

The protected characteristics most commonly cited are race, ethnicity, national origin, religion, sex, gender identity, sexual orientation, and disability. Some frameworks add age, immigration status, caste, or political affiliation. A comparative analysis of hate crime legislation across multiple jurisdictions found substantial variation in which groups receive explicit legal protection, with some countries covering a narrow list and others expanding it over time in response to emerging forms of prejudice.2Enlighten Publications. A Comparative Analysis of Hate Crime Legislation: A Report to the Hate Crime Legislation Review That variation means a statement classified as hate speech in one country may be perfectly legal in another, even when both countries share a general commitment to protecting minority groups.

Why No Single Legal Definition Exists

The post-World War II period fundamentally reshaped how governments thought about hateful expression. Before the war, speech regulation had been primarily about maintaining public order and state authority. The scale of propaganda-driven atrocities during the war revealed what unchecked dehumanizing language could lead to, and the response was a shift toward rights-based regulation. The founding of the United Nations and the drafting of the Universal Declaration of Human Rights established a new framework in which incitement to discrimination and hostility was treated as a human rights concern, not merely a public-order issue.3International Journal of Law Management & Humanities. Historical Evolution of Hate Speech: Pre-World War II to the Digital Era

That rights-based approach took different forms in different legal traditions. Countries like Germany and France enacted criminal prohibitions against Holocaust denial and incitement to hatred, with relatively specific statutory language. The United Kingdom developed a framework of aggravated offenses and public-order restrictions that evolved over decades. The United States, by contrast, has interpreted its First Amendment protections to cover most forms of hateful expression, drawing the line only at “true threats,” incitement to imminent lawless action, and narrowly defined categories of fighting words. This means the same speech act could result in criminal charges in Berlin, platform removal in London, and constitutional protection in New York.

The absence of a unified definition is not just a bureaucratic gap. It reflects genuinely different philosophical commitments about how societies should balance dignity, equality, and free expression. Hate speech regulation is frequently justified as protecting democratic deliberation by ensuring that marginalized groups are not silenced through intimidation. But critics argue that vague or open-ended formulations of what constitutes hate speech can be weaponized to silence political disagreement, creating a chilling effect on legitimate discourse.4International Journal for the Semiotics of Law – Revue internationale de Sémiotique juridique. Hate Speech: Enthymemes, Fallacies and Chilling Effects Both concerns are real, and the tension between them explains why legislatures and courts keep revisiting the question without arriving at a stable answer.

The Difference Between Hate Speech and Offensive Language

One of the most persistent sources of confusion is the boundary between hate speech and language that is simply offensive, vulgar, or rude. A profanity-laced rant about traffic is offensive. A racial slur directed at a specific group with the intent to demean or intimidate crosses into different territory. But many real-world cases fall between those extremes, and the distinction matters for both legal and practical purposes.

Research into automated detection systems has illustrated just how blurry this line can be. One influential study found that classifiers trained on social media data struggled to separate hate speech from other forms of offensive language: racist and homophobic tweets were more likely to be flagged as hate speech, while sexist tweets tended to be classified merely as offensive. Tweets lacking explicit slurs or hate-related keywords were harder to classify correctly regardless of their actual content.5Proceedings of the International AAAI Conference on Web and Social Media. Automated Hate Speech Detection and the Problem of Offensive Language That asymmetry points to a deeper issue: hate speech is not simply a matter of particular words, but of the relationship between the speaker, the target, and the social context in which the words are used.

Researchers studying hate speech frameworks have proposed a tiered approach to help with this problem. Rather than a binary of “hate speech / not hate speech,” one framework classifies messages as allowable, harmful, or inciting. The harmful tier captures speech that demeans or dehumanizes without explicitly calling for action. The inciting tier is reserved for speech that catalyzes or promotes violence or discrimination.6ACM Transactions on Social Computing. Adapting the Dangerous Speech Paradigm to Identify Incitement from Polarized and Hate Speech That three-level system acknowledges something the binary approach misses: the distance between a casually bigoted joke and a deliberate call for ethnic cleansing is enormous, and lumping them together under one label distorts the conversation about what to do about either one.

How Social Media Platforms Define It

For most people, the definition of hate speech that matters most is the one enforced by the platforms where they spend their time. Facebook, X (formerly Twitter), Reddit, YouTube, and TikTok all maintain their own policies, and those policies have evolved considerably over the past two decades. A study examining how Facebook, Twitter, and Reddit defined harassment and hate speech between 2005 and 2020 found that all three platforms moved through distinct phases, with their definitions becoming increasingly complex and nuanced over time.7Policy & Internet. How harassment and hate speech policies have changed over time: Comparing Facebook, Twitter and Reddit (2005–2020)

Early platform policies were either nonexistent or extremely broad, essentially telling users to be civil without specifying what that meant. Over time, platforms introduced lists of protected characteristics, distinguished between direct attacks and general bigotry, and developed escalation pathways for content that was offensive but not clearly in violation. Reddit’s evolution was particularly dramatic: the platform went from a near-absolutist free-speech stance to banning entire communities devoted to hate content, with multiple policy rewrites along the way.

Platform definitions differ from legal definitions in important ways. A platform can ban content that is perfectly legal, and frequently does. It can also fail to remove content that would be illegal in certain jurisdictions. The enforcement is uneven, too. Human moderators reviewing thousands of posts per day inevitably make inconsistent decisions, and automated systems have their own biases, often over-flagging content from marginalized communities while missing subtler forms of hate directed at those same communities.8ResearchGate. Ethical Challenges in AI-Powered Hate Speech Detection: Balancing Automated Moderation and Dataset Creation The training data used to build these systems can inadvertently reinforce stereotypes if the datasets themselves were labeled with biased assumptions about which groups need protection and which do not.

When the Same Word Means Different Things

Context complicates every attempt at definition. The same slur can function as an instrument of oppression when used by an outsider against a marginalized group, or as an act of solidarity and in-group affirmation when reclaimed by members of that group. Detecting reclaimed slurs is one of the hardest challenges for hate speech detection systems, because the lexical item is identical in both cases and only the social identity of the speaker and the conversational context distinguish abuse from affirmation.9CEUR Workshop Proceedings. AIWizards at MULTIPRIDE: A Hierarchical Approach to Slur Reclamation Detection

Coded language adds another layer of difficulty. Dog whistles are expressions that carry a benign surface meaning for the general public and a hostile secondary meaning for a specific audience. The term originated in the context of U.S. political rhetoric, where politicians could signal racial animus to sympathetic listeners without using overtly bigoted language. In recent years, dog whistles have become common on social media as a deliberate strategy for evading content moderation, and they evolve rapidly. Once a coded term is widely recognized and flagged by platforms, users create new ones.10ACL Anthology. Silent Signals, Loud Impact: LLMs for Word-Sense Disambiguation of Coded Dog Whistles This cat-and-mouse dynamic means that any definition of hate speech based primarily on specific words or phrases will always be playing catch-up.

Sarcasm, irony, and humor create similar problems. A sarcastic tweet mocking a racist talking point may use the same language as a sincere endorsement of it. Memes combine images and text in ways that are nearly impossible for automated systems to evaluate without deep cultural knowledge. These edge cases are not rare exceptions. They represent a substantial portion of the content that moderation systems and human reviewers encounter every day.

Real Harm Beyond Hurt Feelings

Debates about hate speech definitions sometimes get framed as a contest between people who are too sensitive and people who value free expression. That framing misses the growing body of evidence on what exposure to hate speech actually does to its targets. A study of online college communities found that users exposed to hateful speech showed measurably higher stress levels compared to matched users who were not exposed. The treatment group’s average stress markers were roughly 39% above their baseline, compared to about 6% for the control group, a statistically significant difference that held even after controlling for offline and online factors.11PubMed Central. Prevalence and Psychological Effects of Hateful Speech in Online College Communities

The effects are not confined to the internet. Research on online hate speech victimization found that victims reported a more pronounced feeling of insecurity outside the internet, not just within it. People targeted by other forms of cybercrime did not show this same spillover, suggesting something particular about hate speech: it attacks a person’s identity in a way that generic online abuse does not, and the resulting anxiety follows them into physical spaces.12Crime Science. Online hate speech victimization: consequences for victims’ feelings of insecurity That finding matters for how we think about definitions. A framework that treats online hate speech as less serious than offline verbal abuse is probably underestimating the harm.

What Works Against It

If defining hate speech is hard, responding to it is harder. The two main approaches are removal (taking content down through moderation or legal enforcement) and counter-speech (responding with messages that challenge the hateful content). Neither has a clear track record of solving the problem.

A systematic review of online interventions designed to reduce hate speech and cyberhate concluded that the evidence is insufficient to determine the effectiveness of these interventions for reducing either the creation or consumption of hateful content.13PubMed Central. Online interventions for reducing hate speech and cyberhate: A systematic review That is a sobering conclusion, but it does not mean nothing works. It means the field has not yet accumulated enough rigorous evidence to draw confident conclusions about which approaches are most effective at scale.

Some individual experiments have shown more promising results. A field experiment on Twitter found that empathy-based counter-speech messages led users to delete xenophobic tweets and reduced the creation of new xenophobic content over the following four weeks. Strategies based on humor or warnings about consequences did not produce consistent effects.14PubMed Central. Empathy-based counterspeech can reduce racist hate speech in a social media field experiment A follow-up line of research found that counter-speech encouraging users to adopt the perspective of minority groups increased tweet deletion by a meaningful amount compared to control groups, while counter-speech based on simple disapproval did not have a significant effect.15Scientific Reports. Counterspeech encouraging users to adopt the perspective of minority groups reduces hate speech and its amplification on social media The pattern across these studies is consistent: appealing to empathy and perspective-taking seems to move the needle in ways that shaming or warning does not.

Whether counter-speech can scale is another question. The experiments involved trained responders crafting individual messages to specific users. Deploying that kind of effort across millions of hateful posts per day is not realistic without some form of automation, which brings us back to the limitations of automated systems. Bots that generate empathetic-sounding counter-speech may lack the authenticity that seems to make such responses effective.

Generative AI as Complication

Large language models have introduced a new dimension to the hate speech problem. These tools can generate persuasive text at scale, and the concern is straightforward: people who want to spread hateful content can use generative AI to produce it quickly, in large volumes, in ways that may be difficult to distinguish from human-written text.16Harvard Data Science Review. The Double-Edged Sword of AI: How Generative Language Models Like Google Bard and ChatGPT Pose a Threat to Countering Hate and Misinformation Online That capability changes the economics of online hate. Previously, producing and distributing hateful propaganda at scale required organized groups and dedicated infrastructure. Now, a single individual with access to widely available tools can generate volumes of content that would have taken teams of human authors to produce.

The same technology is also being deployed on the defensive side. Researchers are training classifiers and language models to detect hate speech with greater sensitivity to context, including experimental work on identifying dog whistles and reclaimed slurs. But the arms race is asymmetric. Generating novel hateful content is computationally cheap, while building systems that reliably identify that content in all its evolving forms requires ongoing investment in training data, annotation, and model refinement. And the definitional problem sits underneath all of it: you cannot train a system to detect something you cannot define, and when the definition shifts across communities, languages, and political contexts, the training data inherits all those disagreements.

The Overbreadth Problem

Any definition broad enough to capture the full range of genuinely hateful expression risks sweeping in speech that most people would consider legitimate, if provocative. This is the overbreadth problem, and it haunts both legal regulation and platform moderation. Hate speech regulation frequently relies on open-ended formulations that invite moral evaluation of the speaker’s views rather than objective assessment of whether harm occurred.17International Journal for the Semiotics of Law – Revue internationale de Sémiotique juridique. Hate Speech: Enthymemes, Fallacies and Chilling Effects When the line between hate speech and strong political opinion depends on a moderator’s or prosecutor’s subjective judgment, the door opens for selective enforcement.

This concern is not hypothetical. Accusations of hate speech have been used strategically in political contexts to discredit opponents, silence dissent, and shut down debate on topics where genuine disagreement exists. At the same time, communities that face real and persistent hateful targeting have legitimate reasons to push for broad definitions and aggressive enforcement. The tension is not between good and bad actors but between two real and competing harms: the harm of unchecked hate speech and the harm of overreaching censorship. Any definition that ignores either side is incomplete.

One practical consequence is that people who belong to marginalized groups sometimes face disproportionate enforcement. Automated moderation systems trained on biased datasets may flag a Black user’s discussion of racism more readily than a white user’s use of racial slurs, because the training data associated certain words with violation regardless of who used them and how. This creates a paradox in which systems designed to protect marginalized communities end up policing their speech more aggressively than the speech directed against them.

Where Evolutionary Psychology Meets Modern Hate

The human tendency to divide the world into in-groups and out-groups is not a product of the internet. Research in evolutionary psychology has described tribalism and parochialism as adaptive responses to ancestral environments in which coalitional aggression between groups was a persistent threat. The inclination to categorize people by group membership and treat in-group members favorably while treating out-group members with suspicion or hostility has deep roots.18PubMed Central. Evolution and the psychology of intergroup conflict: the male warrior hypothesis That does not make modern hate speech natural or inevitable, but it does suggest that the cognitive biases underlying it are not going away through policy alone.

Understanding these roots matters for intervention design. If the impulse to dehumanize out-groups is partly driven by perceived threat, then interventions that reduce perceived threat or increase familiarity with the targeted group should be more effective than interventions based on punishment or shaming. The counter-speech research described earlier aligns with this: empathy and perspective-taking work better than disapproval. Defining hate speech is necessary for legal and platform purposes, but the definition itself does not address the underlying psychology. A comprehensive approach needs both clear rules about what is prohibited and sustained investment in strategies that address why people produce hateful content in the first place.