The modified Rankin Scale is a seven-level grading system that measures how disabled a person is after a stroke, ranging from 0 (no symptoms at all) to 6 (dead). It is the single most common outcome measure in stroke clinical trials and a standard tool in routine neurological care, yet its apparent simplicity masks real tensions: two patients assigned the same grade can have very different abilities, and two clinicians evaluating the same patient sometimes disagree on which grade to assign. Understanding what the scale captures, where it falls short, and how researchers are trying to improve it matters for anyone navigating stroke recovery or reading about stroke treatments.
What Each Grade Means
The scale was originally published by John Rankin in 1957, then modified in the 1980s to include a grade for death so that it could serve as a complete outcome measure in clinical trials. A consensus effort led by the Stroke Therapy Academic Industry Roundtable standardized the language clinicians should use for each level, assigning both a “health state” descriptor and a “valence” term to clarify how good or bad each grade is considered.
1PubMed Central. Standardized Nomenclature for Modified Rankin Scale Global Disability Outcomes: Consensus Recommendations From Stroke Therapy Academic Industry Roundtable XIThe seven grades break down as follows:
- Grade 0: No symptoms at all.
- Grade 1: Symptoms are present but the person is not disabled. They can carry out all usual activities. Considered an “excellent” outcome.
- Grade 2: Slight disability. The person cannot do everything they used to but can look after themselves without help.
- Grade 3: Moderate disability. The person needs some help but can walk without assistance.
- Grade 4: Moderately severe disability. The person cannot walk or attend to bodily needs without help.
- Grade 5: Severe disability. The person is bedridden, incontinent, and requires constant nursing care.
- Grade 6: Dead.
The boundary between grades 2 and 3 carries particular clinical weight. A large prospective study of stroke survivors found a dramatic step-change at that threshold: the adjusted odds of being dead or disabled at five years were roughly 35 times higher for patients discharged at grades 3 through 5 compared with those at grades 0 through 2. Five-year mortality risk nearly doubled across the same divide.
2Stroke. Abstract 166: One-month Modified Rankin Scale (mRS) Score Predicts Five-year Disability, Death, Quality-of-Life, and Healthcare Costs in Ischaemic StrokeHow Reliable Is It Between Different Assessors
One of the persistent concerns about the scale is that different clinicians sometimes assign different grades to the same patient. A systematic review that pooled reliability data across multiple studies found that overall agreement was moderate when measured strictly (unweighted kappa of 0.46) but quite good when close disagreements were given partial credit (weighted kappa of 0.90).
3PubMed. Reliability of the modified Rankin Scale: a systematic reviewWhat that means in practical terms is that assessors rarely disagreed by more than one grade, but exact agreement was only fair. The difference matters: if a trial defines “good outcome” as grade 0 to 2 and your assessor puts you at 2 while another would have said 3, you shift categories entirely.
A structured interview format was developed specifically to tighten this up. In one head-to-head study, pairs of raters evaluating the same patients achieved only 43% exact agreement using the traditional unstructured approach, compared with 81% using the structured interview.
4PubMed. Reliability of the modified Rankin Scale across multiple raters: benefits of a structured interviewA separate study with twelve raters scoring 56 patients found reasonable overall agreement but noted a systematic bias between experienced and inexperienced assessors. Interestingly, the decision tool designed to help inexperienced raters did not improve their performance.
5PubMed. The modified Rankin Scale in acute stroke has good inter-rater-reliability but questionable validityA 2025 systematic review looking at structured validation studies found that weighted agreement between structured and unstructured in-person scores ranged from moderate to very good, with weighted kappa between 0.56 and 0.90. But it also uncovered a consistent pattern at the scale’s extremes: patients with good outcomes tended to rate themselves as doing better on structured questionnaires than their clinicians rated them face-to-face, while patients with poor outcomes rated themselves as doing worse.
6PubMed. Evaluating the strengths and limitations of structured modified rankin scale validation studies – A systematic reviewTelephone and Video Assessment
Following up stroke patients in person at 90 days is expensive and logistically difficult, especially in large multicenter trials. This has driven interest in remote assessment. Telephone-based scoring shows good agreement with face-to-face evaluation, with weighted kappa values across several studies ranging from about 0.71 to 0.89.
7PubMed Central. Reliability of the modified Rankin Scale applied by telephone8Cerebrovascular Diseases. Validation of a Structured Interview for Telephone Assessment of the Modified Rankin Scale in Brazilian Stroke Patients
Video-based assessment appears to do even better. In an endovascular stroke trial, centralized video evaluations achieved a weighted kappa of 0.92 against local face-to-face assessments, compared with 0.77 for centralized phone evaluations.
9PubMed. Phone and Video-Based Modalities of Central Blinded Adjudication of Modified Rankin Scores in an Endovascular Stroke TrialThe advantage of video is not surprising: assessors can observe gait, facial asymmetry, and how the patient interacts with their environment, none of which translates over a phone call. For large clinical trials that need blinded, centralized adjudication, video has become an increasingly popular standard.
The Brazilian Portuguese validation is worth noting because it worked well even in a sample with low educational attainment, suggesting the structured telephone format can function across different literacy levels and cultural contexts.
10Cerebrovascular Diseases. Validation of a Structured Interview for Telephone Assessment of the Modified Rankin Scale in Brazilian Stroke PatientsThe Big Debate in Trials: Where to Draw the Line
For decades, most stroke trials split patients into two groups: “good outcome” (usually grades 0 to 2, sometimes 0 to 1) versus “poor outcome” (everything else). This binary approach is easy to understand and communicate, but it throws away a lot of information. A treatment that shifts many patients from grade 4 to grade 3, for instance, would show no benefit at all in a trial defined by the 0-to-2 cutoff, even though those patients gained meaningful independence.
An alternative called shift analysis, or ordinal analysis, looks at the entire distribution of grades and asks whether the treatment group shifted toward better outcomes across all levels. When researchers reanalyzed data from two landmark trials of clot-dissolving medication (tPA), they found that the shift approach detected a statistically significant benefit in both, including one trial (ECASS-II) that was originally reported as negative using the binary cutoff.
11PubMed. Shift analysis versus dichotomization of the modified Rankin scale outcome scores in the NINDS and ECASS-II trialsThe question of which approach to use is not settled by a single answer, though. Simulation studies show it depends on how the treatment works. For drugs that protect brain tissue broadly (neuroprotective agents), shift analysis is the most efficient way to detect a benefit. For treatments that restore blood flow early, shift analysis and a strict 0-to-1 cutoff perform comparably. For treatments that recanalize vessels late, the traditional 0-to-2 binary cutoff actually works best.
12PubMed Central. Treatment effects for which shift or binary analyses are advantageous in acute stroke trialsThere is also strong evidence that the full ordinal scale does a better job predicting long-term outcomes than either binary split. In a study of more than 1,600 patients, the ordinal mRS was more strongly related to five-year mortality, five-year disability, and five-year care costs than either the 0-to-1 or 0-to-2 dichotomy. The ordinal model reduced prediction error for care costs by about $3,000 per patient compared with an age-and-sex model alone, versus roughly $2,800 for the 0-to-2 split and only $1,600 for the 0-to-1 split.
13PubMed Central. Ordinal vs dichotomous analyses of modified Rankin Scale, 5-year outcome, and cost of strokeMajor recent trials, including those that proved the benefit of mechanical clot retrieval (thrombectomy), have used the mRS as their primary endpoint. The SWIFT PRIME trial, for example, defined its primary outcome as 90-day global disability assessed by the scale.
14PubMed Central. Solitaireâ„¢ with the Intention for Thrombectomy as Primary Endovascular Treatment for Acute Ischemic Stroke (SWIFT PRIME) trial: protocol for a randomized, controlled, multicenter study comparing the Solitaire revascularization device with IV tPA with IV tPA alone in acute ischemic strokeWhat the Scale Misses
The mRS captures broad functional independence, but it was never designed to reflect the full picture of recovery. A study of 73 patients with upper-extremity weakness after stroke illustrated this starkly: within grade 2 alone, upper-extremity impairment ranged from near-complete paralysis to no measurable deficit at all. A substantial number of patients experienced meaningful improvements in arm function and other measures over 90 days yet still ended up at grade 3 or worse, appearing to have a “poor outcome” by conventional trial definitions even though they had recovered considerably.
15PubMed Central. Association of Modified Rankin Scale With Recovery Phenotypes in Patients With Upper Extremity Weakness After StrokeThe scale also does not capture cognitive outcomes well. A study comparing the mRS to a frailty index for predicting neurocognitive disorders three months after stroke found that the frailty index was a stronger predictor. When both were entered into the same statistical model, the mRS’s predictive strength dropped more than the frailty index’s did, suggesting that whatever the frailty index captures about pre-stroke vulnerability adds information the mRS alone does not provide.
16PubMed Central. Is Frailty Index a better predictor than pre-stroke modified Rankin Scale for neurocognitive outcomes 3-months post-stroke?Compared with the NIHSS, a neurological examination scale that scores specific deficits like vision, speech, and limb strength, the mRS is also less sensitive in detecting treatment effects. A comparison across acute stroke trials found that the NIHSS allowed smaller sample sizes or greater statistical power because it picks up finer neurological changes that the mRS’s broad categories can miss.
17PubMed. Comparison of the National Institutes of Health Stroke Scale with disability outcome measures in acute stroke trialsTranslating Grades Into Quality of Life
Health economists need a way to convert mRS grades into quality-of-life values for cost-effectiveness analyses. A large pooled study of nearly 23,000 acute stroke patients mapped mRS scores onto a standard quality-of-life instrument (the EQ-5D) and derived utility weights for each grade. The values tell a clear story: grade 0 corresponded to a utility of 0.96 (essentially full health), grade 1 to 0.83, grade 2 to 0.72, grade 3 to 0.54, grade 4 to 0.22, and grade 5 to negative 0.18, meaning that patients at that level of disability rated their health state as worse than death on average. Grade 6 was assigned 0 by definition.
18PubMed. Utility-Weighted Modified Rankin Scale Scores for the Assessment of Stroke Outcome: Pooled Analysis of 20 000+ PatientsThese utility-weighted scores let researchers combine survival and disability into a single number, which is increasingly used as an alternative to the traditional binary endpoint. Rather than asking “did the patient reach grade 0 to 2 or not,” a utility-weighted analysis captures the fact that moving from grade 4 to grade 3 represents a much larger quality-of-life gain than moving from grade 1 to grade 0. Mapping algorithms have also been developed to translate mRS scores into other utility instruments, giving health economists flexibility depending on which quality-of-life measure a given country or health system prefers.
19PubMed. Mapping the modified Rankin scale (mRS) measurement into the generic EuroQol (EQ-5D) health outcome20PubMed. Mapping the modified rankin scale (mRS) onto the assessment of quality of life (AQoL) utilities
Use Beyond Ischemic Stroke
Although the mRS was built for stroke, it has been adopted as an outcome measure in other conditions where brain injury causes lasting disability. In subarachnoid hemorrhage (bleeding around the brain, often from a ruptured aneurysm), the scale is routinely used at discharge. One study found that admission severity strongly predicted mRS outcome: patients classified as intermediate severity had about six times the odds of a poor discharge score compared with low-severity patients, and high-severity patients had roughly 66 times the odds.
21Emory University OpenEmory. Association between Hunt and Hess Grade and the Modified Rankin Scale among Patients with Non-Traumatic Subarachnoid HemorrhageWhether the scale works well for subarachnoid hemorrhage survivors is another question. An international survey found poor agreement between patients’ self-assessed outcomes, their actual mRS scores, and the conventional binary split into “good” versus “poor.” The authors argued that patient-centered measurement tools are needed for this population rather than simply borrowing the stroke scale wholesale.
22PubMed Central. Patient Relevance of the Modified Rankin Scale in Subarachnoid Hemorrhage: An International Cross-sectional SurveyThe scale also appears in traumatic brain injury research. In a study of patients with minor head injuries, the pre-injury mRS score was a strong predictor of both traumatic intracranial hemorrhage and short-term morbidity at discharge.
23PubMed. Risk of Intracranial Hemorrhage and Short-Term Outcome in Patients with Minor Head InjuryIn pediatric stroke, the mRS has been compared with scales designed specifically for children. One study found a very strong correlation between the mRS and the Pediatric Stroke Outcome Measure overall, but the correlation was much weaker for language and cognitive subscales, reinforcing that the mRS captures physical disability better than it captures thinking and communication problems.
Caregiver Burden and the Downstream Costs of Each Grade
The scale has implications well beyond clinical trial endpoints. For every grade on the mRS, the demands placed on caregivers escalate. A Nigerian study of stroke survivors and their caregivers found that caring for patients with worse mRS scores was associated with roughly four-fold increases in caregiver burden, and caring for incontinent survivors drove burden even higher.
24PubMed Central. Predictors of caregiver burden after stroke in Nigeria: Effect on psychosocial well-beingHealthcare costs follow a similar gradient. The ordinal analysis study mentioned earlier showed that each step up the mRS is associated with higher five-year care costs, and that the full ordinal scale captures cost differences better than any binary cutoff. This matters for health systems trying to decide where to invest: if a treatment shifts a meaningful number of patients down by even one grade, the downstream savings in care and lost productivity can be substantial, even when the shift does not cross the traditional “good outcome” threshold.
25PubMed Central. Ordinal vs dichotomous analyses of modified Rankin Scale, 5-year outcome, and cost of strokeAutomating the Score With AI
Assigning an mRS grade traditionally requires a trained assessor talking to the patient or reviewing their records. With millions of stroke admissions generating clinical notes every year, researchers have started exploring whether artificial intelligence can extract mRS scores directly from electronic health records. A fine-tuned large language model trained on clinical notes achieved 92% accuracy when sorting patients into the binary good-versus-poor categories (grades 0 to 2 versus 3 to 6), and 77% accuracy when predicting all seven individual grades. Its weighted kappa of 0.92 for the multiclass task suggests that when the model got it wrong, it was typically off by only one grade.
26PubMed Central. Assessment of the Modified Rankin Scale in Electronic Health Records With a Fine-Tuned Large Language Model: Development and Internal ValidationA separate effort built a natural language processing pipeline to pull mRS scores from unstructured clinical notes automatically, aiming to support large-scale observational research where manually scoring every patient is impractical.
27PubMed Central. Automated extraction of post-stroke functional outcomes from unstructured electronic health recordsThese tools are still in the validation phase and not yet replacing human assessors in trials. But they point toward a future where mRS data can be harvested at scale from routine clinical documentation, opening up stroke outcome research to datasets that were previously too large and messy to score by hand. The irony is that the scale’s simplicity, the very quality that makes it vulnerable to criticism for being too coarse, is exactly what makes it tractable for machine learning: seven clearly ordered categories map onto a classification problem that current AI handles well.

