The PHQ, or Patient Health Questionnaire, is a family of brief self-report screening tools used to flag depression, anxiety, and physical symptoms in everyday medical settings. Developed in the late 1990s as a self-administered version of an older clinician-led diagnostic instrument called PRIME-MD, the PHQ was designed so that patients could fill it out themselves in a waiting room without needing a trained interviewer. The most widely used version, the PHQ-9, asks nine questions about depressive symptoms over the past two weeks, and it has become one of the most common mental health screening instruments on the planet. But the PHQ-9 is only one member of a broader toolkit that includes ultra-short screeners, anxiety measures, and a somatic symptom scale, each suited to different clinical situations.
Where the PHQ Came From
Before the PHQ existed, primary care doctors who wanted to screen for mental health conditions had to use the PRIME-MD, which required a clinician to conduct a structured interview. That took time most doctors didn’t have. In 1999, researchers published validation data showing that a self-administered paper questionnaire could match the diagnostic accuracy of the clinician-led version while being far more practical for busy clinics.1PubMed. Validation and utility of a self-report version of PRIME-MD: the PHQ primary care study That questionnaire became the PHQ. Its appeal was immediate: it was free to use, required no special training to administer, and could be completed in a few minutes. Those properties helped it spread rapidly through primary care worldwide.
The PHQ Family at a Glance
People often say “the PHQ” when they mean the PHQ-9, but the full family includes several instruments designed for different purposes:
- PHQ-9: Nine items measuring depression severity over the past two weeks. Each item is scored 0 to 3, giving a total range of 0 to 27. This is the flagship tool.
- PHQ-2: Just the first two items of the PHQ-9, covering depressed mood and loss of interest. Used as an initial quick screen to decide whether the full PHQ-9 is warranted.
- PHQ-4: Combines the PHQ-2 with two anxiety items from the GAD-2, creating a four-question screener that flags both depression and anxiety simultaneously.
- PHQ-15: Fifteen items asking about common physical symptoms like headaches, stomach pain, and fatigue. Used to gauge the burden of somatic symptoms, which often accompany depression or exist on their own.
Each version trades off between brevity and depth. A doctor’s office that screens every patient at every visit might start with the PHQ-2 and only move to the full PHQ-9 if those first two questions raise a flag. A research study tracking treatment outcomes might use the PHQ-9 at every visit for finer detail.
How Accurately Does the PHQ-9 Detect Depression
The standard threshold for a “positive” screen on the PHQ-9 is a score of 10 or above. A large individual-participant meta-analysis found that at this cutoff, sensitivity was about 88% and specificity about 85% when tested against semistructured diagnostic interviews.2BMJ. Accuracy of Patient Health Questionnaire-9 (PHQ-9) for screening to detect major depression: individual participant data meta-analysis In practical terms, that means the PHQ-9 catches the large majority of people who truly have major depression while correctly ruling out most of those who don’t.
There has been debate about whether the cutoff should be a point or two lower. A separate meta-analysis pooling data across many studies found that cutoff scores anywhere between 8 and 11 showed acceptable diagnostic properties, with no sharp cliff between them.3PubMed Central. Optimal cut-off score for diagnosing depression with the Patient Health Questionnaire (PHQ-9): a meta-analysis A more recent large analysis found that lowering the cutoff to 8 yielded roughly 80% sensitivity and 82% specificity across a broad, unweighted population.4JAMA Network Open. Data-Driven Cutoff Selection for the Patient Health Questionnaire-9 Depression Screening Tool The upshot is that the standard cutoff of 10 works well in most settings, but clinicians sometimes adjust it a few points depending on whether they’d rather catch more cases at the cost of more false positives or vice versa.
The Two-Question Shortcut
The PHQ-2 strips the questionnaire down to its essentials: “Over the last two weeks, how often have you been bothered by little interest or pleasure in doing things?” and “How often have you been bothered by feeling down, depressed, or hopeless?” A meta-analysis of studies using semistructured interviews as the reference standard found that at a cutoff of 2 or higher, the PHQ-2 had about 91% sensitivity and 67% specificity, while raising the cutoff to 3 or higher dropped sensitivity to about 72% but improved specificity to around 85%.5JAMA. Accuracy of the PHQ-2 Alone and In Combination With the PHQ-9 for Screening to Detect Major Depression: Systematic Review and Meta-analysis
Another diagnostic meta-analysis concluded that the PHQ-2 is adequate as a first-step screen in primary care, with good acceptability from patients, though the cutoff threshold needs to be chosen carefully depending on the clinical goal.6PubMed Central. Case finding and screening clinical utility of the Patient Health Questionnaire (PHQ-9 and PHQ-2) for depression in primary care: a diagnostic meta-analysis of 40 studies The standard workflow in many clinics is to administer the PHQ-2 first and then follow up with the full PHQ-9 only if the short version flags a concern. This keeps the screening burden low for the majority of patients who are not depressed.
Screening for Depression and Anxiety at Once
Depression and anxiety frequently overlap. The PHQ-4 addresses this by combining the two depression items from the PHQ-2 with two anxiety items from the GAD-2 (the abbreviated Generalized Anxiety Disorder scale). Validation studies confirm that the PHQ-4 holds up as a two-factor tool, with one factor capturing depression and the other anxiety, and it maintains structural consistency across different age and gender groups.7PubMed. A 4-item measure of depression and anxiety: validation and standardization of the Patient Health Questionnaire-4 (PHQ-4) in the general population A more recent U.S. study found high internal consistency for the overall scale and its subscales, and confirmed the two-factor structure, though it also showed that the depression and anxiety factors were very strongly correlated with each other.8PubMed Central. Reliability and validity of the Patient Health Questionnaire-4 scale and its subscales of depression and anxiety among US adults based on nativity
In routine clinical practice, the PHQ-9 and the GAD-7 (the full seven-item anxiety scale) are often administered together. One real-world analysis of concurrent scores found that the two scales are highly correlated, with more than half of paired scores falling into the same severity category.9PubMed. Value added? A pragmatic analysis of the routine use of PHQ-9 and GAD-7 scales in primary care That tight correlation has prompted some researchers to question whether using both adds much clinical information beyond what one alone provides, though the two scales do diverge in individual cases in ways that can guide treatment choices.
The Suicide Risk Question
Item 9 of the PHQ-9 asks how often the person has been bothered by “thoughts that you would be better off dead, or of hurting yourself.” This single item has drawn substantial research attention because it sits at the intersection of screening and safety. A study of outpatients found that one-year cumulative risk of a suicide attempt rose from about 0.4% among those who answered “not at all” to roughly 4% among those who answered “nearly every day.” The risk of death by suicide over one year rose tenfold across the same range, from about 0.03% to 0.3%.10PubMed Central. Does response on the PHQ-9 Depression Questionnaire predict subsequent suicide attempt or suicide death?
A separate analysis looking across age groups found that people reporting nearly daily suicidal ideation on item 9 were five to eight times more likely to attempt suicide within 30 days compared to those without such thoughts, and the elevated risk persisted over two years.11PubMed Central. Suicidal Ideation Reported on the PHQ9 and Risk of Suicidal Behavior across Age Groups Item 9 was never designed to be a standalone suicide risk assessment tool, but these findings confirm that any positive response warrants follow-up.
Using the PHQ-9 to Track Treatment
Beyond initial screening, the PHQ-9 is widely used to monitor how well treatment is working. The standard definitions of treatment “response” and “remission” come from the instrument itself: response is usually defined as a 50% reduction in total score, and remission as a score dropping below 5. In practice, these benchmarks vary somewhat across health systems. One large analysis found that three-month response rates ranged from about 32% to 51% depending on which metric was applied, with remission around 22% in that same window.12PubMed Central. Assessing the Impact of Different Depression Treatment Success Metrics on Organizational Performance Longer follow-up periods consistently show higher success rates, reflecting the reality that depression recovery often takes months.
The PHQ-9 tracks treatment change about as well as the BDI-II, a more established depression scale used mainly in research and specialty settings.13PubMed. Psychometric comparison of the PHQ-9 and BDI-II for measuring response during treatment of depression It also correlates strongly with clinician-rated scales like the Hamilton Depression Rating Scale, with correlation coefficients in the range of 0.6 to 0.7 across studies.14PubMed Central. The Patient Health Questionnaire-9 vs. the Hamilton Rating Scale for Depression in Assessing Major Depressive Disorder15PubMed Central. The reliability and validity of PHQ-9 in patients with major depressive disorder in psychiatric hospital That said, one study noted that using the PHQ-9 as a pure outcome measure may be less sensitive to change than dedicated rating scales, especially in research contexts where fine-grained distinctions matter.16PubMed. Defining successful treatment outcome in depression using the PHQ-9: a comparison of methods
Physical Symptoms and the PHQ-15
Many patients present to their doctor with physical complaints that have no clear medical explanation, or that overlap heavily with depression and anxiety. The PHQ-15 was developed to quantify the severity of such somatic symptoms. It asks about stomach pain, back pain, headaches, dizziness, chest pain, and other common complaints. Scores of 5, 10, and 15 mark thresholds for low, medium, and high somatic symptom severity, and higher scores predict worse functioning, more sick days, and more healthcare visits in a clear stepwise pattern.17PubMed. The PHQ-15: validity of a new measure for evaluating the severity of somatic symptoms
A large systematic review and meta-analysis of the PHQ-15’s measurement properties confirmed good internal consistency overall but identified a few weak spots: items about menstrual problems, fainting spells, and sexual problems correlated poorly with the rest of the scale.18JAMA Network Open. Measurement Properties of the Patient Health Questionnaire–15 and Somatic Symptom Scale–8: A Systematic Review and Meta-Analysis Those items may not tap the same underlying construct as the others, which is worth knowing if you’re interpreting a patient’s total score. The PHQ-15 also shows convergent validity with depression and general distress scales, reinforcing the tight connection between unexplained physical symptoms and psychological distress.19Psychosomatics. Psychometric Properties of the Patient Health Questionnaire–15 (PHQ–15) for Measuring the Somatic Symptoms of Psychiatric Outpatients
Special Populations
Pregnant and Postpartum Women
Perinatal depression is common and often underdiagnosed. A systematic review and meta-analysis found that the PHQ-9, using the standard cutoff of 10, performed well during pregnancy and the postpartum period, with pooled sensitivity of 84% and specificity of 81%. Those numbers were virtually identical to those of the Edinburgh Postnatal Depression Scale (EPDS), the screening instrument traditionally used in maternity settings.20PubMed Central. Screening for Perinatal Depression with the Patient Health Questionnaire Depression Scale (PHQ-9): A Systematic Review and Meta-analysis This means clinics that already use the PHQ-9 for their general patient population don’t necessarily need a separate screening instrument for perinatal care.
Adolescents and Children
A modified version of the PHQ-9, called the PHQ-9M or PHQ-9A, has been adapted for younger populations. In adolescent psychiatric clinic patients, the modified version showed strong internal consistency and good convergent validity with other depression rating scales, though it was somewhat less sensitive to treatment-related changes in symptom severity than clinician-administered scales.21PubMed Central. Psychometric Properties of the Patient Health Questionnaire-9 Modified for Major Depressive Disorder in Adolescents More recently, researchers tested the PHQ-9A in children as young as 10 and found that at an optimal cutoff of 5, sensitivity was 93% and specificity was 60% for detecting major depressive disorder, and kids above that cutoff who didn’t meet full diagnostic criteria still showed significant impairment.22PubMed. Validity of the PHQ-9A as a self-report screener for major depressive disorder in youth ages 10-12 years That lower cutoff reflects the reality that children tend to endorse fewer symptoms on self-report measures, so the bar needs to be set lower to catch them.
Older Adults
The PHQ-9 works in elderly primary care populations, but there is a wrinkle. One study found that its diagnostic accuracy was highest in older patients with fewer than three medical comorbidities, where the area under the curve reached 0.93. In patients burdened with multiple chronic conditions, accuracy dipped, likely because symptoms like fatigue and poor appetite overlap between depression and medical illness.23PubMed Central. A study of the diagnostic accuracy of the PHQ-9 in primary care elderly Still, another study specifically examining elderly patients with diabetes and chronic lung disease concluded that PHQ-9 sum scores remained a valid and reliable screening method even in those populations.24PubMed. Summed score of the Patient Health Questionnaire-9 was a reliable and valid method for depression screening in chronically ill elderly patients The takeaway for older adults is that the PHQ-9 still works, but a positive screen deserves especially careful clinical follow-up because physical illness can inflate scores.
Cross-Cultural and Language Considerations
The PHQ-9 has been translated into dozens of languages, and a recurring research question is whether scores mean the same thing across cultural and linguistic groups. Studies have tested this with statistical methods that look for items behaving differently in different groups. Comparing German- and Turkish-language versions, researchers found that individual items like “sleep problems,” “appetite changes,” and “anhedonia” showed some differential functioning across groups, but total scores remained unbiased.25PubMed Central. Cross-cultural validation of the German and Turkish versions of the PHQ-9: an IRT approach A study comparing English and French versions of the PHQ-9 found the same pattern: three of nine items showed statistically detectable differences, but the overall depression score was not substantively affected.26PLoS ONE. Are Scores on English and French Versions of the PHQ-9 Comparable? An Assessment of Differential Item Functioning
Testing in a Nepali-speaking primary care population showed that the standard cutoff of 10 performed well, with 94% sensitivity and 80% specificity against a structured clinical interview.27PubMed Central. Detection of depression in low resource settings: validation of the Patient Health Questionnaire (PHQ-9) and cultural concepts of distress in Nepal Research on American Indian and Alaska Native populations found that the PHQ-9 maintained adequate measurement equivalence when compared to diverse racial and ethnic groups, supporting its use across these populations.28PubMed Central. Evaluating the Cross-Cultural Measurement Invariance of the PHQ-9 between American Indian/Alaska Native Adults and Diverse Racial and Ethnic Groups The consistent finding is reassuring: individual items may behave a bit differently across cultures, but the total score appears to measure the same underlying thing regardless of language version.
Digital and App-Based Administration
With the rise of telehealth and mobile health apps, a natural question is whether filling out the PHQ-9 on a phone or tablet gives the same result as doing it on paper. A study using data from a large community-dwelling population found that the small difference in scores between app-based and traditional administration was not clinically meaningful, and among participants with major depression, there was no difference at all.29Value in Health. Differences in Patient Health Questionnaire-9 Scores by Mode of Administration: Findings from the Project Baseline Health Study A pilot study using tablets in primary care waiting rooms found that screening rates were higher when tablets were used compared to the usual paper-based workflow, though staff responsible for distributing the devices had concerns about scalability.30PubMed Central. A pilot study of participatory and rapid implementation approaches to increase depression screening in primary care
That said, one comparison study cautioned that the transition from paper to electronic formats is not always seamless and may require proper adaptation rather than simply digitizing the same questions.31Revista de PsiquiatrÃa y Salud Mental (English Edition). Comparative study of pencil-and-paper and electronic formats of GHQ-12, WHO-5 and PHQ-9 questionnaires In practice, most validated digital implementations of the PHQ-9 reproduce scores closely enough for clinical use, and the convenience advantages are substantial.
Known Limitations and False Positives
The PHQ-9 is a screening tool, not a diagnostic test. Screening tools are built to cast a wide net, which means they inevitably pull in people who don’t actually have the condition. A study in primary care patients with diabetes or heart disease found that applying the PHQ-9’s diagnostic algorithm resulted in high false positive rates, meaning a large percentage of people were incorrectly flagged as having minor or major depression.32Journal of Affective Disorders. Diagnostic accuracy of the Patient Health Questionnaire-9 for assessment of depression in type II diabetes mellitus and/or coronary heart disease in primary care This is particularly relevant in medically ill populations where symptoms like fatigue, poor sleep, and appetite changes could reflect the medical condition rather than depression. A high PHQ-9 score should always be the start of a conversation with a clinician, not the end of it.
A related issue is that how you instruct someone to interpret the questions matters more than you might expect. A recent randomized trial found that when participants were given specific instructions emphasizing the distinction between anxiety symptoms and depression symptoms, their scores dropped substantially compared to a control group given standard instructions. The effect was large, roughly 2.6 points on both the PHQ-9 and the GAD-7.33JAMA Network Open. Instructional Emphasis and PHQ-9 and GAD-7 Scores Among Participants With Anxiety or Depression: A Randomized Clinical Trial That finding suggests the scores people give are influenced by how they frame their own symptoms, which adds a layer of noise to any single administration.
Implementing Screening in Real-World Clinics
Having a valid screening tool is one thing; getting it used consistently is another. Quality improvement projects have demonstrated that implementation efforts can dramatically increase screening rates. One pediatric primary care initiative took adolescent depression screening from 0% to nearly 75% and saw clinically meaningful increases in diagnoses, mental health referrals, and treatment.34PubMed. Implementation of Universal Adolescent Depression Screening: Quality Improvement Outcomes A similar project focused on adolescent screening found that educating staff about the screening tool and establishing a workflow led to measurable improvements in screening rates.35PubMed. Implementation of an Evidence-Based Clinical Guideline for Depression Screening of the Adolescent The recurring theme is that the PHQ tools themselves are ready; the bottleneck is integrating them into clinical workflows so they actually get used.
Is Universal Screening Worth the Cost
Economic evaluations suggest that systematic PHQ-based depression screening is a good investment. A modeling study of PHQ screening combined with collaborative care in New York City found that starting universal screening at age 20 would cost about $660 per person over a 50-year horizon and gain roughly 0.38 quality-adjusted life years per person, yielding a cost-effectiveness ratio of about $1,726 per quality-adjusted life year. That figure is far below conventional thresholds for what health systems consider worth spending.36PubMed Central. The cost-effectiveness of PHQ screening and collaborative care for depression in New York City An Indian economic evaluation reached an even more favorable conclusion, finding that universal PHQ-9 screening in primary care was actually cost-saving from a societal perspective once indirect costs like lost productivity were factored in.37The Lancet Regional Health – Southeast Asia. Economic impact and cost-utility of systematic PHQ-9 depression screening in primary health care in India: a model-based economic evaluation These analyses strengthen the case that was already building from clinical evidence: screening with the PHQ-9 doesn’t just identify depression, it sets in motion a chain of care that pays for itself.

