Content Analysis in Research: Methods, Coding, and LLMs

Content analysis is a research method for systematically examining communication, whether written text, images, audio, or video, by applying consistent coding rules to identify patterns, themes, or frequencies in the material. The technique dates back to theological studies in the late 1600s, when the Church scrutinized nonreligious printed materials it viewed as threatening, and has since expanded into virtually every field that studies how people communicate.1SAGE Publications. Content Analysis: An Introduction to Its Methodology – History – Section: History What seems like straightforward “read and categorize” work actually involves careful decisions about what to count, how to count it, and how to demonstrate those counts are trustworthy. Those decisions have gotten considerably more complicated with the arrival of social media, multilingual datasets, and artificial intelligence.

Quantitative and Qualitative Approaches

Content analysis splits broadly into two camps: quantitative and qualitative. The quantitative version converts communication into numbers. You define categories ahead of time, apply a structured coding scheme, and use statistical methods to analyze the results. The emphasis is on replicability: another researcher following the same codebook should arrive at the same numbers. This approach works well for testing a specific hypothesis, such as whether news coverage of a political candidate skews negative or whether corporate press releases mention environmental responsibility more frequently than they did a decade ago.

Qualitative content analysis, by contrast, cares more about meaning than counting. Researchers read through material looking for themes, sometimes guided by prior theory (a deductive approach) and sometimes letting patterns emerge from the data itself (an inductive approach). Inductive analysis tends to be used when there is little existing research on the topic, while deductive analysis works better for testing whether established frameworks hold up in a new setting or time period.2PubMed. The qualitative content analysis process The trade-off is real: quantitative approaches sacrifice nuance for consistency, while qualitative approaches sacrifice consistency for depth. Many studies now blend the two, using qualitative coding to generate categories and then quantifying how often those categories appear.3Standardisierte Inhaltsanalyse in der Kommunikationswissenschaft – Standardized Content Analysis in Communication Research. Content analysis in mixed method approaches

What Gets Coded and Why It Matters

One of the first decisions in any content analysis is whether you are looking at manifest content or latent content. Manifest content is what is explicitly present on the surface: specific words, phrases, images, or actions that can be directly observed and counted. Latent content is the underlying meaning, tone, or intent behind what is communicated. Both approaches are recognized in the methodological literature as valid but serving different purposes.4PubMed Central. Demystifying Content Analysis

Counting how many times a news article mentions “climate change” is manifest analysis. Deciding whether the article frames climate change as an urgent crisis or a distant concern is latent analysis. The manifest approach is easier to make reliable because two coders can agree on whether a word appears. The latent approach captures richer meaning but depends more on the coder’s interpretation, which introduces disagreement. A recent study testing whether AI could handle both types found that automated tools performed well on manifest coding, picking out and categorizing surface-level content with high accuracy, but struggled with the interpretive nuance required for latent analysis.5PubMed Central. Evaluating Microsoft Copilot in Qualitative Health Research: Accurate for Manifest Content Coding but Limited in Latent Interpretation That gap matters because much of what researchers want to understand about communication lives in the latent layer.

Making Sure the Coding Is Trustworthy

A content analysis is only as good as the agreement between the people (or machines) doing the coding. If two coders read the same news story and one codes it as “positive toward the candidate” while the other codes it as “neutral,” the study has a problem. This is why inter-rater reliability is treated as a gate that a study must pass before its findings are taken seriously.

Several statistical measures exist for assessing this agreement, including variants of the kappa statistic, intraclass correlation, and Krippendorff’s alpha.6PubMed Central. Some common statistical methods for assessing rater agreement in radiological studies Krippendorff’s alpha is widely used because it handles different numbers of coders, different types of measurement scales, and missing data. The commonly accepted thresholds are that a score of 0.80 or above indicates reliable agreement, scores between 0.67 and 0.79 warrant caution, and anything below 0.67 suggests the coding is too inconsistent to draw firm conclusions from.7MethodsX. K-Alpha Calculator–Krippendorff’s Alpha Calculator: A user-friendly tool for computing Krippendorff’s Alpha inter-rater reliability coefficient A score of zero means the coders agreed no more than you would expect by random chance, and negative scores mean they systematically disagreed, which usually signals a broken codebook rather than just sloppy work.

Frameworks like Gwet’s agreement coefficient have also been developed to extend to situations with many raters and many rating categories.8The Stata Journal: Promoting communications on statistics and Stata. Implementing a General Framework for Assessing Interrater Agreement in Stata The choice of which measure to report depends on the study design, but the underlying question is always the same: would someone else coding this material reach the same conclusions?

Automated Tools and Their Limits

The volume of digital communication has made purely manual content analysis impractical for many research questions. If you want to study how millions of tweets discussed a political event, you cannot have graduate students read each one. This pressure has driven the development of automated tools, and one of the most widely used has been LIWC (Linguistic Inquiry and Word Count), a software program that counts words belonging to predefined psychological and linguistic categories. LIWC has been adapted for text analysis in multiple languages, including a Serbian adaptation that showed acceptable equivalence with the English version.9Psihologija. Psychometric evaluation of the Serbian dictionary for automatic text analysis – LIWCser

But automated dictionary tools have a well-documented weakness: they operate at the word level and miss context. A study evaluating LIWC’s reliability on large online datasets found that precision dropped to as low as about 50% and recall to about 42% for some categories, meaning the tool was miscategorizing roughly half the content in certain domains.10Applied Corpus Linguistics. Is LIWC reliable, efficient, and effective for the analysis of large online datasets in forensic and security contexts? You could manually correct the errors, but that defeats the purpose of using automated analysis on a large scale. Custom dictionaries built for specific research questions, like a dictionary designed to capture self-transcendent emotions in text, can perform better within their narrow domain, though they still struggle with rhetorical devices like sarcasm and metaphor.11PLoS ONE. Developing and validating the self-transcendent emotion dictionary for text analysis

Topic modeling represents a different automated approach. Rather than counting predefined categories, algorithms like Latent Dirichlet Allocation (LDA) identify clusters of words that tend to appear together, allowing researchers to discover themes they did not specify in advance. Semi-supervised versions of these models let researchers seed the algorithm with some initial guidance, combining human insight with computational scale.12Social Science Computer Review. Integrating Human Insights Into Text Analysis: Semi-Supervised Topic Modeling of Emerging Food-Technology Businesses’ Brand Communication on Social Media These methods are useful for exploratory work but still require careful human evaluation of the topics they produce.

Large Language Models as Coders

The most significant recent disruption to content analysis has been large language models. Researchers have started using LLMs not just as writing tools but as stand-ins for human coders. A study comparing LLM performance to outsourced human coders on complex text classification tasks found that all tested LLMs outperformed the human coders across every task, with the best-performing models only marginally beating each other.13Scientific Reports. LLMs outperform outsourced human coders on complex textual analysis The study also observed that as models grow more powerful, the gap between them and outsourced coders widens rather than closes.

Practical guides have emerged for researchers who want to use LLMs to augment or replace human coders during content analysis, including approaches that replicate original study codings, scale analysis to cover ten times as many articles, and extend to additional languages.14Social Science Computer Review. A Practical Guide and Case Study on How to Instruct LLMs for Automated Coding During Content Analysis The appeal is obvious: LLMs can process thousands of documents in hours, work across languages, and do not get tired or bored. But important caveats remain. The comparison in most studies is between LLMs and outsourced, minimally trained coders, not expert researchers who have spent years developing deep familiarity with the material. LLMs also tend to perform well on manifest coding tasks but lose ground on latent interpretation that requires contextual judgment.15PubMed Central. Evaluating Microsoft Copilot in Qualitative Health Research: Accurate for Manifest Content Coding but Limited in Latent Interpretation The field is moving fast, and the practical question for most researchers is no longer whether to use LLMs at all but which tasks they can handle reliably and which still need human eyes.

Where Content Analysis Gets Applied

The method’s flexibility is part of why it shows up almost everywhere. Political communication research has relied on it for decades. Studies have used content analysis to examine how news media frame political candidates, including how the framing differs across political systems. A multinational study of online news spanning nearly 2,700 stories found that media in majoritarian democracies like the United States produced more candidate-centered, simplistic, and polarizing character framing than media in multiparty consensus democracies.16The International Journal of Press/Politics. Multinational and Multimodal Character Framing of Political Candidates in Online News: Do Political and Media System Classifications Matter? Content analysis of broadcast news has been used to examine how networks frame the press itself and the publicity process surrounding campaigns.17American Behavioral Scientist. Framing the Press and the Publicity Process

Health communication is another major area. A content analysis of online patient-provider communication in China identified 15 distinct communication strategies that doctors used, and found that specific strategies like providing clear information, making diagnoses, expressing empathy, and enabling patients to manage their own care were linked to higher patient satisfaction.18PubMed Central. How do provider communication strategies predict online patient satisfaction? A content analysis of online patient-provider communication transcripts In these applied settings, the value of content analysis is that it turns unstructured communication into evidence that can inform training, policy, or platform design.

Beyond Text: Analyzing Images and Video

Content analysis is no longer limited to words on a page. As communication has become increasingly visual, researchers have developed methods to systematically analyze images, video, and multimedia content. Visual models for social media analysis now include approaches for grouping images by theme, tracking which images generate the most engagement, following visual trends over time, and comparing how images rank across different platforms.19International Journal of Communication. Visual Models for Social Media Image Analysis: Groupings, Engagement, Trends, and Rankings

Computer vision algorithms have made large-scale visual content analysis feasible. A study of major European corporations analyzed over 21,000 website images and more than 3,600 visual Twitter posts to examine how businesses visually communicate about sustainability, using automated content analysis to compare what companies show versus what a balanced approach to social, environmental, and economic responsibility would look like.20Corporate Social Responsibility and Environmental Management. Visualizing the triple bottom line: A large‐scale automated visual content analysis of European corporations’ website and social media images The multinational political framing study mentioned earlier also coded across six modalities, including still images, moving images, frozen video frames, text, audio, and superimposed text, reflecting how modern content analysis increasingly treats communication as multimodal rather than text-only.21The International Journal of Press/Politics. Multinational and Multimodal Character Framing of Political Candidates in Online News: Do Political and Media System Classifications Matter?

Sampling Social Media at Scale

One practical challenge that has grown alongside digital communication is sampling. Social media platforms generate enormous volumes of content daily, and deciding which posts, comments, or images to include in a study is far from straightforward. Researchers have explored various sampling strategies and found that different sample sizes and methods can produce similar findings when managed carefully, but that working with massive corpora like collections of tens of millions of tweets requires deliberate choices about extraction and storage before analysis even begins.22PubMed Central. Strategies for the Analysis of Large Social Media Corpora: Sampling and Keyword Extraction Methods

The difficulty is compounded by platform-specific quirks. Application programming interfaces (APIs) that researchers use to collect data may return non-random samples, rate-limit how much data can be pulled, or change access rules unexpectedly. A dataset collected through keyword searches will miss relevant posts that use different terminology, while a dataset collected through user sampling will miss relevant posts by users not in the sample. None of these problems are unique to content analysis, but they interact with it directly because the validity of a content analysis depends on having a defensible sample to code.

Cross-Cultural Complications

Running a content analysis across languages and cultures introduces a set of problems that do not exist in single-language studies. A coding category that works cleanly in English may not have a direct equivalent in Japanese or Korean. A gesture or visual cue that reads as friendly in one culture may be neutral or even rude in another. Research on cross-cultural content analysis of television advertising across Japan, Korea, and the United States has emphasized that achieving reliable results requires deep understanding of the relevant languages and cultures, consistent selection and training of coders, and careful back-translation of research instruments to ensure equivalence.23Emerald Insight. Achieving reliable and valid cross-cultural research results in content analysis

This is an area where automated tools can both help and hurt. Machine translation makes it possible to process content in languages the researcher does not speak, and LLMs can code across languages, but the interpretive subtleties that matter most in cross-cultural work are precisely the ones that automated tools handle least well. A coding prompt that works for English-language news articles may produce systematically different results when applied to a German-language corpus, even if the underlying material covers the same topic. Researchers extending LLM-based coding to additional languages have found it necessary to validate performance separately for each language rather than assuming that accuracy transfers.24Social Science Computer Review. A Practical Guide and Case Study on How to Instruct LLMs for Automated Coding During Content Analysis

Ethics of Analyzing Online Content

When content analysis focused primarily on published news articles, books, or government documents, ethical questions were relatively simple: the material was public and intended for broad audiences. The shift toward analyzing social media posts, forum discussions, and user-generated content has made things murkier. Key ethical challenges center on whether online content should be treated as public or private, whether researchers need informed consent from the people whose posts they analyze, and what anonymity and privacy look like in a digital landscape where supposedly anonymized quotes can be traced back to their authors with a quick search.25Emerald Insight. Reframing Qualitative Research Ethics

There is no universal consensus. A tweet posted to a public account is technically public, but the person who wrote it may not have anticipated it being scraped into a dataset and coded by researchers. Health-related posts in support communities occupy an especially sensitive zone. Most institutional review boards now require researchers to consider these questions, but the answers vary by institution, platform, and country. The general trend has been toward treating user-generated content with more caution than traditional public documents, even when it is technically accessible to anyone.

Open Science and the Reproducibility Gap

A persistent weakness in published content analyses is that other researchers often cannot fully reproduce them. The codebook, the raw data, and the analytic code are the three pieces someone would need to replicate a study, and all three are frequently missing from published work. A content analysis of articles published in 2025 across 20 leading communication journals evaluated how often researchers shared their materials, data, and code. The results were discouraging: while research materials like codebooks were shared more frequently than in the past, data and analytic code were rarely made available, and overall transparency remained inconsistent.26Media and Communication. The Adoption of Open Science Practices in Communication Research: Taking Stock of the Field in 2025

Some researchers have proposed solutions. One approach uses minimally but systematically trained online workers to code content, with reliability achieved by aggregating the judgments of multiple coders. This setup matches the reliability and validity of traditional intensively trained research assistants but offers much greater speed and, crucially, transparency and replicability, because the codebook and instructions can be shared openly without needing to also share the specialized training that took place in a lab.27Political Science Research and Methods. Online coders, open codebooks: New opportunities for content analysis of political communication LLM-based coding pushes this further: if the prompt and model version are documented, anyone can rerun the analysis and check whether they get the same results. The infrastructure for genuinely reproducible content analysis is now more accessible than it has ever been. Whether the field will actually use it is a separate question.