Most surveys in the United States draw their race categories from a federal standard maintained by the Office of Management and Budget (OMB). For decades, the minimum set included five racial groups and a separate ethnicity question about Hispanic or Latino origin. In 2024, the OMB revised those standards for the first time in nearly thirty years, merging race and ethnicity into a single question and adding a Middle Eastern or North African category. The result is that the “list of races” on any given survey depends on when it was designed, who designed it, and what country it serves, and those choices shape who gets counted accurately and who falls through the cracks.
The Standard Federal Categories
Since 1997, the OMB required all federal agencies and any survey tied to federal data collection to use at least five racial categories: White, Black or African American, American Indian or Alaska Native, Asian, and Native Hawaiian or Other Pacific Islander. Ethnicity was handled separately with just two options: Hispanic or Latino, and Not Hispanic or Latino. These categories were never intended to be biological or scientific classifications. They were administrative groupings created to track civil-rights compliance, health disparities, and population trends. Race, as researchers have consistently described it, is a complex, multidimensional construct shaped by social and political forces, not a clean biological taxonomy.1PubMed Central. Conceptualizing and categorizing race and ethnicity in health services research
The 2024 revision, known as SPD 15, made several significant changes. It combined the race and ethnicity questions into a single item, added Middle Eastern or North African (MENA) as a standalone category, and required a write-in field for detailed responses. The updated minimum categories are now: American Indian or Alaska Native, Asian, Black or African American, Hispanic or Latino, Middle Eastern or North African, Native Hawaiian or Pacific Islander, and White. Federal agencies have until 2029 to implement the changes, so you will encounter both the old and new formats for years to come.
Why Race and Ethnicity Were Separate in the First Place
Under the old OMB standard, a respondent answered two questions: first, “Are you Hispanic or Latino?” and second, “What is your race?” This two-question format treated Hispanic/Latino identity as an ethnicity that could overlap with any race. In theory, a person could identify as Hispanic and White, or Hispanic and Black, or Hispanic and any other racial category. In practice, the setup confused many respondents. A study of high school students found that roughly 92% of students ended up classified the same way regardless of whether they answered a single combined question or the two-question format, but the two-question version produced more missing data, particularly among Black students who left later questions blank.2PubMed. High school student responses to different question formats assessing race/ethnicity
Research on the old NIH measure, which followed the two-question OMB structure, found it demonstrated poor differentiation between race and ethnicity, offered restricted response options, and lacked an inclusive ethnicity question. Separating the two concepts and giving respondents more flexibility to identify themselves both racially and ethnically was recommended as a way to improve validity.3PubMed Central. “Which box should I check?”: examining standard check box approaches to measuring race and ethnicity The 2024 OMB revision effectively responded to this critique by collapsing the two questions into one and expanding the options.
The Hispanic and Latino Classification Problem
One of the most persistent tensions in race survey design has been where Hispanic and Latino identity fits. Under the old system, Hispanic was not a race. A person could check “Hispanic or Latino” for ethnicity and then had to pick a separate racial box. Many Hispanic respondents found this confusing or unsatisfying, since they understood their identity as encompassing both racial and ethnic dimensions simultaneously. Research using open-ended questions alongside the standard census-style format has found substantial misalignment between how Hispanic respondents classify themselves on a closed-ended race question and how they describe their racial identity in their own words. Factors like self-rated skin tone and whether someone is a first-generation or later-generation immigrant strongly influenced these choices.4Socius: Sociological Research for a Dynamic World. Counting Race, Miscounting Identity: The Census and Hispanic Racial Self-Identification
The practical consequence was that a huge share of Hispanic respondents chose “Some Other Race” on the census, making it one of the largest racial categories in the country despite being essentially a catch-all. This undermined the data quality that the classification system was supposed to provide. The 2024 revision, by making Hispanic or Latino a category alongside the racial options on a single question, aims to resolve this long-standing mismatch between how the government categorizes people and how people actually see themselves.
Middle Eastern and North African Respondents
For decades, people of Middle Eastern or North African descent were officially classified as White on federal forms. There was no separate box. Research has consistently shown that this classification does not match how MENA individuals experience their own identity or how others perceive them. A study published in the Proceedings of the National Academy of Sciences found that both non-MENA White Americans and MENA Americans considered MENA-related traits, including ancestry, names, and religion, to be MENA rather than White. When given the option, most MENA individuals chose to identify as MENA or as both MENA and White, with second-generation individuals and those who identify as Muslim particularly likely to embrace a distinct MENA identity.5PubMed Central. Middle Eastern and North African Americans may not be perceived, nor perceive themselves, to be White
The same research found that MENA Americans who perceive more anti-MENA discrimination are more likely to identify with a MENA label, suggesting that the experience of racial hostility activates a stronger group identity. The authors argued that as long as MENA Americans remained aggregated with Whites, any inequalities they face would stay hidden in the data. The 2024 OMB revision directly addressed this by adding MENA as its own category. If you are designing a survey and working from the new federal standard, MENA should appear as a distinct option.
Multiracial Identification and “Check All That Apply”
Before the year 2000, the US Census required respondents to choose a single race. Starting with the 2000 Census, a “mark one or more” instruction replaced the single-selection format. This was a major shift, but its effects depend heavily on how the question is designed and how the data get processed downstream.
A laboratory study of multiracial and multiethnic women found that respondents’ racial identification varied considerably depending on the question format. People of mixed heritage preferred formats that gave them the opportunity to acknowledge their multiracial background rather than forcing them into a single box.6Evaluation Review. Dimensions of Self-Identification Among Multiracial and Multiethnic Respondents in Survey Interviews Research using census ancestry data has also explored how people with multiracial backgrounds several generations back identify themselves, finding that the population who might reasonably claim multiple racial ancestries is much broader than just children of two single-race parents.7Social Science Research. Choosing race: Multiracial ancestry and identification
The practical challenge for survey designers is what happens after data collection. If your analysis requires putting each person into a single racial group, you need a decision rule for how to handle people who checked multiple boxes. Some researchers assign a primary race based on algorithms; others create a separate “multiracial” category. Either approach loses information. If your survey goals allow it, keeping the multi-select data intact gives a more accurate picture, but it complicates comparisons with older datasets that forced a single choice.
American Indian and Alaska Native Data
American Indian and Alaska Native (AI/AN) populations face a specific and well-documented problem with survey race categories: they get collapsed into an “Other” bucket. When sample sizes are small, researchers and reporting agencies frequently merge AI/AN data with Asian, Native Hawaiian, and Pacific Islander respondents into a single residual category. This practice erases meaningful health and social differences between these groups. Researchers have argued that collapsing these categories into “Other” provides no benefit to public health policymakers, researchers, or tribal planners, and that tribal affiliation should be collected whenever feasible.8PubMed Central. Office of Management and Budget racial categories and implications for American Indians and Alaska Natives
A related problem is misclassification. An analysis of the Youth Risk Behavior Surveillance System found that how AI/AN youth are classified can differ significantly depending on whether a survey uses self-reported race directly or aggregates it using algorithms. Misclassification, noncollection, or the use of categories such as “Other” and “multirace” without allowing disaggregation can distort estimates of disease burden and health outcomes for AI/AN populations.9PubMed Central. Comparing Self-Reported and Aggregated Racial Classification for American Indian/Alaska Native Youths in YRBSS: 2021 If you are building a survey that might reach AI/AN respondents, the takeaway is to keep the category distinct, include it in your reporting, and consider adding a tribal affiliation field.
Open-Ended Versus Closed-Ended Questions
One alternative to checkbox lists is simply asking people to describe their race or ethnicity in their own words. A study testing an open-ended collection system in a health care setting found that agreement between the open-ended responses and standard closed categories was about 93%, with a high statistical concordance. Rates of missing data and the proportion of people categorized as “Other” were actually lower with the open-ended approach. Latino/Hispanic and multiracial/multiethnic individuals were more likely to prefer using their own words to describe their identity.10PubMed Central. A system for rapidly and accurately collecting patients’ race and ethnicity
The downside of open-ended responses is the work required to code them afterward. Free-text answers produce hundreds of unique strings that need to be mapped back to standard categories for analysis. For small-scale research or clinical intake, that coding burden is manageable. For large national surveys processing millions of responses, it can become a bottleneck. The emerging middle ground is to offer both: a set of standard checkboxes plus a write-in field for detail. This is the approach the 2024 OMB revision now requires for federal data collection.
How Survey Mode Changes Responses
Whether a survey is administered in person, on paper, by phone, or on the web can affect how people report their race. A study comparing in-person and web-administered screening found that race had somewhat lower consistency between modes than other demographic characteristics. While the overall match was still about 87%, it was the lowest among all characteristics measured. Most of the discrepancies occurred between White and Other, with no clear direction suggesting that one mode systematically pushed people toward a particular category.11PLoS ONE. Comparability of in-person and web screening: Does mode affect what households report?
A separate study looking at medical records found significant differences between the way race was recorded in electronic health records and how patients described themselves, particularly among those who identified as Hispanic. The mismatch was concentrated in the White and Other categories, echoing the pattern seen in survey mode comparisons.12PubMed Central. Discrepancies in Race and Ethnicity Documentation: a Potential Barrier in Identifying Racial and Ethnic Disparities If you are collecting race data across multiple channels, such as both online and in-person registration, expect some inconsistency and plan for reconciliation.
How Other Countries Handle It
The American approach of listing explicit racial categories is not universal. Different countries handle race, ethnicity, and ancestry in strikingly different ways depending on their histories and political frameworks.
France, for example, has historically refused to collect racial or ethnic data in its national census at all. The French republican tradition treats citizens as individuals rather than members of racial groups, and collecting such data has been viewed as antithetical to universalism. The emergence of “diversity” as a public concern since the 2000s has renewed debate about whether racial and ethnic data should be integrated into the census, but the tension between tracking discrimination and upholding colorblind universalism remains unresolved.13European Journal of Cultural Studies. The uses of universalism. ‘Diversity Statistics’ and the race issue in contemporary France
Brazil takes yet another approach. Its census uses five color/race categories (branco, preto, pardo, amarelo, and indÃgena), and researchers have studied how these compare to more granular self-assessment tools like a 1-to-10 skin color scale. A comparative study found that both methods capture essentially the same underlying construct, but that Brazilians tend to place themselves toward the center of the color gradient, and women classify themselves slightly lighter than men on the scale, an effect not seen with the standard categories.14PubMed Central. Comparison between two race/skin color classifications in relation to health-related outcomes in Brazil The Brazilian system reflects a different cultural understanding of race as a spectrum rather than a set of discrete groups.
The United Kingdom asks about “ethnic group” rather than “race,” with categories like White British, Black Caribbean, Pakistani, Bangladeshi, and Mixed. Canada collects data on “visible minorities” and Indigenous identity using categories tied to the country’s employment equity legislation. Each system reveals its society’s particular history with race and colonialism, and none maps cleanly onto the American categories. If you are designing a survey for an international audience, picking one country’s race list and applying it globally will produce confused respondents and unreliable data.
Practical Tips for Building a Race Question
If you are writing a survey and need to include a race question, several evidence-based principles can improve your data quality:
- Follow the 2024 OMB standard if federal compliance matters. That means a single combined race-and-ethnicity question with at least seven categories (American Indian or Alaska Native, Asian, Black or African American, Hispanic or Latino, Middle Eastern or North African, Native Hawaiian or Pacific Islander, and White) plus a write-in option. Allow respondents to select more than one.
- Add sub-categories where your research needs them. “Asian” encompasses people of Chinese, Indian, Filipino, Vietnamese, Korean, Japanese, and many other ancestries whose health profiles and socioeconomic conditions differ. If your survey goals require that granularity, add it. The same goes for “Black or African American,” which includes African Americans, Caribbean Americans, and recent immigrants from African countries.
- Keep AI/AN as a distinct category. Do not collapse it into “Other” during analysis. If sample sizes are small, acknowledge the limitation rather than erasing the population from your results.
- Include a write-in field. Even if you plan to map free-text responses back to standard categories, giving respondents the chance to describe themselves increases satisfaction and captures information that checkboxes miss.
- Use self-identification, not observer assignment. Having a clerk or interviewer guess a person’s race introduces systematic errors, particularly for Hispanic, multiracial, and MENA respondents.
Reducing Missing and Unknown Data
One of the most common problems with race data is not which categories appear on the form but how many responses come back blank or coded as unknown. A case study from a large urban health system found that a dedicated data-capture improvement program reduced unknown race and ethnicity fields by about 76% over five years, from 454 unidentified patients in 2016 to 107 in 2020.15PubMed Central. Improving Patient Race and Ethnicity Data Capture to Address Health Disparities: A Case Study From a Large Urban Health System The improvement came from training front-line staff, standardizing the collection process, and making race and ethnicity fields required rather than optional in the electronic system.
The lesson generalizes beyond health care. Whatever your survey context, the biggest threat to your race data is usually not the wrong categories but missing responses. Clear instructions, a visually clean layout, self-identification rather than third-party assignment, and treating the field as important rather than optional all help. If your survey is online, making the race question required (with a “Prefer not to answer” option for those who genuinely decline) dramatically reduces blanks without forcing people to choose an identity they reject.
When the Categories Do Not Fit
Even the best-designed list will leave some respondents feeling unseen. Afro-Latino individuals may feel torn between Black and Hispanic. People from Central Asian countries may not see themselves in “Asian” as typically understood in the US. Someone whose family is from Guyana might identify as Indo-Caribbean and find no box that captures that. Pacific Islanders from Melanesia have pointed out that their experiences differ significantly from those of Polynesian populations, yet both are grouped under a single heading.
No finite list of checkboxes can capture the full range of human racial and ethnic identity. The categories on any survey are tools for measurement, not definitions of who people are. The goal is not perfection but adequacy: can the data you collect answer the questions you are asking, without systematically erasing or miscounting any group? If your survey is for a specific community or purpose, tailoring the options to that context almost always beats copying a generic federal list and calling it done.

