Propensity refers to a natural tendency or inclination toward something, and the word shows up across remarkably different fields with distinct technical meanings in each. In statistics and medical research, a “propensity score” is a specific mathematical tool used to approximate randomized experiments when true randomization is impossible. In criminology, “criminal propensity” describes a stable tendency toward offending, often tied to self-control. In genetics, propensity captures the inherited likelihood of developing a disease. These uses share a common thread, the idea that some underlying disposition shapes outcomes, but the methods researchers use to measure and apply that disposition vary enormously.
The Propensity Score as a Statistical Tool
The most technically influential use of “propensity” comes from a 1983 paper by Paul Rosenbaum and Donald Rubin, which defined the propensity score as the probability that someone receives a particular treatment given their observed characteristics.1Biometrika. The central role of the propensity score in observational studies for causal effects The insight was elegant: if you can estimate this single number for every person in your study, adjusting for it removes bias from all the observed characteristics at once. That paper has become one of the most cited in all of statistics, and propensity score methods now show up constantly in medical research, public health, economics, and the social sciences.2PubMed Central. Bias associated with using the estimated propensity score as a regression covariate
The reason this tool matters so much is practical. Randomized controlled trials, where researchers flip a coin to decide who gets a treatment, are the gold standard for determining cause and effect. But you cannot always randomize. You cannot randomly assign people to smoke or not smoke, to receive one drug versus another when clinical judgment intervenes, or to participate in a job-training program versus staying home. In these observational settings, the people who get treatment differ systematically from those who do not. Propensity scores offer a way to account for those differences.
Matching, Weighting, and Other Approaches
Once you have estimated each person’s propensity score, there are several ways to use it. The two most common are matching and weighting, and they work quite differently in practice.
Propensity score matching pairs each treated person with an untreated person who has a similar score. Researchers have developed a range of algorithms for doing this, including optimal matching, nearest-neighbor matching with or without replacement, and caliper-based matching that only allows pairs within a specified score distance.3PubMed Central. A comparison of 12 algorithms for matching on the propensity score The choice of algorithm is not trivial. Monte Carlo simulations comparing twelve different matching strategies found that the details, such as whether you match from highest to lowest score or in random order, and whether you allow a treated person to match with only one or multiple untreated persons, can meaningfully affect results.4PubMed Central. A comparison of 12 algorithms for matching on the propensity score When you have many confounders, finding well-matched pairs can become difficult or even impossible; the propensity score helps by collapsing all those confounders into a single dimension.5Journal of Biostatistics and Epidemiology. Comparison of Nearest Neighbor and Caliper Algorithms in Outcome Propensity Score Matching to Study the Relationship between Type 2 Diabetes and Coronary Artery Disease
Inverse probability of treatment weighting, or IPTW, takes a different approach. Instead of discarding unmatched individuals, it reweights everyone in the sample so that the treated and untreated groups look more alike. This tends to retain more of the original data, which increases the effective sample size compared with matching.6PubMed Central. An introduction to inverse probability of treatment weighting in observational research IPTW is also more straightforward to implement because it avoids the thorny decisions about which matching algorithm to use and what to do with unmatched participants.7JAMIA Open. Accurate treatment effect estimation using inverse probability of treatment weighting with deep learning That said, weighting introduces its own problems, especially when some individuals have extreme propensity scores close to zero or one. Those individuals receive enormous weights that can destabilize the entire estimate.
Checking Whether It Actually Worked
A propensity score analysis is only as good as the balance it achieves. After matching or weighting, researchers need to verify that the treated and untreated groups actually look similar on baseline characteristics. The most widely used check is the standardized mean difference, which measures how far apart the two groups are on each characteristic using a common scale. Because it does not depend on units of measurement, you can compare balance across variables as different as age, blood pressure, and income on the same chart.8PubMed Central. Balance diagnostics after propensity score matching
But standardized mean differences only capture averages. Researchers also check the ratio of variances between groups, compare higher-order moments and interactions, and use graphical tools like quantile-quantile plots and side-by-side boxplots to spot remaining imbalances.9PubMed Central. Balance diagnostics for comparing the distribution of baseline covariates between treatment groups in propensity-score matched samples A study that reports propensity score matching without these diagnostics is essentially asking you to trust, without evidence, that the method worked. The field has gotten better about this over time, but published papers still sometimes skip the step.
Where Propensity Scores Can Fail
Propensity scores are powerful, but they rest on assumptions that do not always hold. The most common failures fall into a few categories.
The first is unmeasured confounding. Propensity scores only adjust for characteristics that the researcher actually observed and included in the model. If there is some important variable that differs between the treated and untreated groups but was never measured, the propensity score cannot fix that bias. This is the fundamental limitation of all observational methods, and it is why researchers have developed sensitivity analyses that estimate how strong an unmeasured confounder would have to be to change the study’s conclusions.10PubMed Central. Using Sensitivity Analyses for Unobserved Confounding to Address Covariate Measurement Error in Propensity Score Methods
The second is violations of the positivity assumption. This is the requirement that, within every combination of characteristics, some people receive treatment and some do not. When certain types of patients always or never receive a treatment, like when a drug is contraindicated for a particular group, the propensity scores for those patients get pushed toward zero or one.11PubMed Central. Core concepts in pharmacoepidemiology: Violations of the positivity assumption in the causal analysis of observational data: Consequences and statistical approaches In weighting analyses, these extreme scores produce a small number of enormously influential weights that can lead to biased estimates and wild variability in results.12PubMed. Propensity score weighting under limited overlap and model misspecification
Perhaps the most counterintuitive failure is what researchers call the propensity score matching paradox. As matching becomes more aggressive, pruning pairs with the largest score differences in pursuit of closer matches, it can actually increase imbalance on individual covariates, make the results more dependent on the specific model used, and introduce more bias rather than less.13PubMed Central. Implications of the Propensity Score Matching Paradox in Pharmacoepidemiology This happens because the propensity score is a summary of many variables, and two people can have nearly identical scores for completely different reasons. Matching them closely on the score does not guarantee they are similar on the underlying characteristics that matter.14PubMed Central. Propensity Score Matching: should we use it in designing observational studies? This finding has led some methodologists to argue that matching should always be accompanied by thorough balance checks and that researchers should think carefully about whether weighting or stratification might be more appropriate.
Machine Learning Meets Propensity Scores
Traditionally, propensity scores are estimated using logistic regression, a workhorse statistical model. But as datasets have grown larger and more complex, researchers have experimented with machine learning methods. Boosting algorithms, random forests, and neural networks can all estimate the probability of treatment assignment, and they are often better at capturing complicated nonlinear relationships between characteristics.15PubMed Central. Propensity score estimation: machine learning and classification methods as alternatives to logistic regression
The evidence on whether this actually helps is mixed, though. A recent benchmarking study compared machine learning-based propensity scores against traditional logistic regression in a real-world heart failure dataset, then checked both against the results of a randomized trial. The logistic regression model using expert-selected confounders came closest to the trial result. The machine learning approach, specifically generalized boosting models combined with data-driven variable selection, did not outperform the traditional method and appeared to amplify bias through overadjustment.16medRxiv. Evaluation of Machine Learning-Based Propensity Score Estimation: A Benchmarking Observational Analysis Against a Randomized Trial The takeaway is that fancier algorithms are not automatically better. Subject-matter knowledge about which variables actually confound the treatment-outcome relationship still matters more than computational sophistication.
In labor economics, the picture is slightly more nuanced. One study evaluating job-training programs for the long-term unemployed found that LASSO-based logistic models improved propensity score estimates in smaller, high-dimensional datasets, while random forests sometimes made things worse when the share of treated individuals was low. In larger samples with higher treatment rates, the choice of estimation method mattered less.17Labour Economics. Does the estimation of the propensity score by machine learning improve matching estimation? The case of Germany’s programmes for long term unemployed
Propensity Scores Beyond Medicine
While pharmacoepidemiology, the study of drug effects in real-world populations, has been one of the biggest consumers of propensity score methods,18PubMed Central. Propensity Scores in Pharmacoepidemiology: Beyond the Horizon the technique has spread far beyond clinical research. In labor economics, propensity score matching is described as the “major workhorse” for evaluating job-training and employment programs.19Labour Economics. Does the estimation of the propensity score by machine learning improve matching estimation? The case of Germany’s programmes for long term unemployed Evaluations of active labor market programs in Serbia, for example, used propensity score matching to reduce the bias that comes from systematic differences between people who participated in employment programs and those who did not, finding positive effects on employment probability.20Economic Annals. The Use Of Propensity Score-Matching Methods In Evaluation Of Active Labour Market Programs In Serbia
Survey research has also adopted the concept. When pollsters collect data from nonprobability sources like social media panels, the people who respond differ systematically from the general population. Propensity score weighting can adjust for this by estimating each respondent’s probability of being in the nonprobability sample versus a traditional probability sample, then reweighting accordingly. One demonstration of this approach reconciled social media survey data with a probability sample on 25 out of 27 political attitudes assessed.21PubMed Central. A Demonstration of Propensity-Score Weighting to Adjust a Social Media Nonprobability Sample Survey of Political Attitudes As traditional polling becomes more expensive and response rates decline, this kind of statistical correction is gaining importance.
Time-varying treatments present another frontier. Standard propensity scores assume that treatment is assigned at a single point in time, but many real-world exposures change over the course of a study. Someone might start a medication, stop it, switch to another, and restart. Time-dependent propensity scores attempt to handle this, though their usefulness depends on how much the changing treatment itself influences future confounders.22PubMed Central. Performance of time-dependent propensity scores: a pharmacoepidemiology case study
Criminal Propensity and Self-Control Theory
Outside of statistics, “propensity” has a long history in criminology. Gottfredson and Hirschi’s General Theory of Crime, published in 1990, argued that crime can be explained primarily by a single factor: low self-control, which they treated as a stable propensity established early in life. Their claim was bold and deliberately provocative. They argued that this criminal propensity does not change much over time, that past offending has no causal effect on future offending once you account for this underlying trait, and that the apparent influence of life events like marriage or employment on crime is largely spurious.23Criminology. MULTIPLE ROUTES TO DELINQUENCY? A TEST OF DEVELOPMENTAL AND GENERAL THEORIES OF CRIME
A meta-analysis of the empirical evidence found that low self-control is indeed a meaningful predictor of crime and related behaviors, regardless of how researchers measured it.24Criminology. THE EMPIRICAL STATUS OF GOTTFREDSON AND HIRSCHI’S GENERAL THEORY OF CRIME: A META‐ANALYSIS That said, the stronger claims of the theory, that propensity is essentially fixed and that no other factors matter, have not fared as well. Developmental theories argue that there are multiple routes to delinquency and that life circumstances genuinely alter trajectories, not just correlate with an unchanging disposition. The debate remains active, but most criminologists now treat self-control as one important risk factor among several rather than the single explanation Gottfredson and Hirschi proposed.
Genetic Propensity and Polygenic Risk
When people talk about having a genetic “propensity” for a disease, they usually mean that inherited variants across the genome collectively raise or lower risk. Researchers capture this idea with polygenic risk scores, which sum up the small contributions of thousands of genetic variants into a single number reflecting someone’s genetic predisposition. Recent work has pushed the accuracy of these scores considerably. In one large study, combining the outputs of multiple scoring algorithms and adding in clinical characteristics like age, sex, and known risk factors pushed twelve out of thirty disease models past 80% accuracy as measured by the area under the curve.25PubMed Central. Optimization of multi-ancestry polygenic risk score disease prediction models
A persistent challenge is that most genetic studies have been conducted in populations of European ancestry, and scores derived from those studies predict less well in other groups. Newer ensemble methods that integrate data from multiple ancestries are narrowing this gap. One approach called PRSmix+ improved prediction accuracy by roughly 70% in European populations and about 40% in South Asian populations compared with standard single-method scores.26Cell Genomics. Comprehensive evaluation and integration of polygenic risk scores across diverse populations Getting these scores to work equitably across populations is one of the major ongoing challenges in genomic medicine.
An important nuance is that genetic propensity rarely operates in isolation. Researchers studying gene-environment interactions, specifically whether polygenic scores for conditions like ADHD or depression interact with environmental risks like harsh discipline or family instability, have found that any such interactions tend to be small. In one study that tested 48 possible gene-environment combinations, the largest significant interaction explained only about 0.4% of the variance in conduct problems.27PubMed Central. Gene-environment interaction using polygenic scores: Do polygenic scores for psychopathology moderate predictions from environmental risk to behavior problems? Genetic propensity is real, but the popular notion that a bad environment “activates” risky genes in some dramatic way overstates what the data typically show.
Behavioral Propensity in Neuroscience
In neuroscience, propensity often comes up in the context of risk-taking and decision-making. Why do some individuals consistently choose the risky option, whether in gambling tasks, substance use, or everyday choices? Research using animal models has identified circuits connecting the prefrontal cortex, the amygdala, the striatum, and dopamine-producing brain regions that together regulate how organisms weigh certain rewards against uncertain ones.28PubMed Central. Neural mechanisms regulating different forms of risk-related decision-making: Insights from animal models Dysfunction in these circuits shows up in a range of psychiatric conditions, from addiction to pathological gambling, suggesting that an individual’s propensity for risky choices is not just a personality quirk but reflects measurable differences in brain wiring.
Animal behavior research has extended the propensity concept even further. Individual animals within the same species show consistent behavioral differences across time and situations, what researchers now call “animal personalities.” Some individuals are consistently bolder, more exploratory, or more aggressive. An interesting and underexplored factor is parasitism: because bolder or more social animals may face higher parasite exposure, and because parasites alter the host’s physical condition, local parasite environments may actually shape the evolution of behavioral propensities in wild populations.29PubMed Central. Parasitism and the evolutionary ecology of animal personality The same disposition that makes an animal a better forager might also make it a more attractive target for parasites, creating evolutionary trade-offs that maintain variation in behavioral propensity within a population.
Nudging and the Propensity to Choose
Behavioral economics uses “propensity” in a looser but practically significant way: the tendency of people to choose one option over another depending on how the choice is presented. Choice architecture, the design of the environment in which decisions are made, exploits these propensities without restricting freedom. A meta-analysis spanning multiple behavioral domains found that nudges, interventions like changing default options, reordering items on a menu, or simplifying enrollment forms, reliably shift behavior by leveraging people’s existing propensities toward the path of least resistance.30PubMed Central. The effectiveness of nudging: A meta-analysis of choice architecture interventions across behavioral domains The propensity here is not a fixed trait so much as a predictable response pattern. People are not choosing retirement savings or organ donation based on deep deliberation in most cases; they are going with whatever feels easiest, and that tendency is remarkably consistent across cultures and domains.
This connects back, indirectly, to the statistical concept. Propensity scores work because people’s likelihood of receiving treatment is predictable from their characteristics. Nudges work because people’s likelihood of choosing an option is predictable from the choice environment. In both cases, the useful insight is the same: propensity, whether toward a treatment, a behavior, or a disease, is not random. It follows patterns. And once you can estimate those patterns, you can either adjust for them in research or deliberately shape them in policy.

