The F1 score is a single number that summarizes how well a classifier performs by combining two more basic metrics, precision and recall, into their harmonic mean. It ranges from 0 (total failure) to 1 (perfect precision and perfect recall), and it has become one of the most commonly reported evaluation metrics in machine learning and natural language processing. But the F1 score carries assumptions and blind spots that can mislead you if you don’t understand what it actually measures and, just as importantly, what it ignores.
What Precision and Recall Actually Tell You
To understand why the F1 score exists, you need the two ingredients it blends. Precision answers: of all the items the model flagged as positive, how many were actually positive? A spam filter that rarely flags legitimate mail as spam has high precision. Recall answers: of all the items that truly were positive, how many did the model catch? A spam filter that catches every spam message has high recall. These two goals pull against each other. You can achieve perfect recall by flagging everything as spam, but your precision drops to nearly zero. You can achieve perfect precision by flagging nothing, but then recall is zero. Almost every real-world model sits somewhere in the tension between these two.
The F1 score resolves that tension into one number by taking the harmonic mean of precision and recall. Unlike a regular average, the harmonic mean penalizes extreme imbalances. If precision is 0.95 but recall is 0.10, a simple average would give you a reassuring-sounding 0.525. The harmonic mean gives you roughly 0.18, which much more honestly reflects how badly the model is failing on one front. That built-in penalty for lopsidedness is the core reason the F1 score caught on so widely, especially in fields like information retrieval and text classification where missing relevant results is a serious failure.
Why It Took Over in Machine Learning
The F1 score originated in information retrieval, where researchers needed a way to judge whether a search engine was finding the right documents. It gained momentum because plain accuracy, the percentage of all predictions that are correct, falls apart when the classes are imbalanced. If only 1% of emails are spam, a model that labels everything “not spam” achieves 99% accuracy while catching zero spam. F1 avoids that trap because it focuses specifically on how the model handles the positive class, measuring whether the model finds positives and whether the things it calls positive are correct.
This sensitivity to the positive class made F1 especially popular in tasks like named-entity recognition, medical diagnosis, fraud detection, and sentiment analysis. In natural language processing, it is routinely treated as the core criterion for judging model performance.1Scientific Reports. Nested named entity recognition in traditional Chinese medicine electronic medical records via dual-granularity feature augmentation and span classification Its use has expanded far beyond text retrieval into nearly every corner of supervised classification, though that expansion has also drawn criticism, which we’ll get to shortly.
Micro, Macro, and Weighted Variants
When you have more than two classes, the F1 score needs to be adapted, and how you adapt it matters a lot. The two most common approaches are micro-averaged F1 and macro-averaged F1, and they can give you very different pictures of the same model.
Micro-averaged F1 pools all the predictions together across every class, then computes precision and recall from the totals. This means large classes dominate the result. If your model handles a class with 10,000 examples well but completely fails on a class with 50 examples, micro-averaged F1 will look fine because the big class drowns out the small one. Macro-averaged F1 computes F1 separately for each class, then takes the unweighted average. This gives every class equal voice, so a failure on a rare class drags the score down just as much as a failure on a common one.2SpringerLink (Appl Intell). Confidence interval for micro-averaged F1 and macro-averaged F1 scores
There is also a weighted-average F1, which computes per-class F1 scores and then averages them weighted by the number of examples in each class. This sits somewhere between micro and macro. The choice among these variants is not cosmetic. In a medical setting where a rare disease class matters enormously, macro-averaged F1 is usually more appropriate. In a high-volume document classification system where most classes are roughly equal in size, micro-averaged F1 works fine. Reporting which variant you used is essential, yet many published papers just say “F1” without specifying, which makes their results difficult to compare.
The Positive-Class Problem
One of the most important things to understand about F1 is that it is asymmetric. It measures how well you handle the positive class and completely ignores how well you handle the negative class. True negatives, the cases the model correctly identifies as negative, play no role whatsoever in the F1 calculation. This design made sense in the original information-retrieval setting, where the “negative class” was the vast ocean of irrelevant documents and correctly ignoring them was not very interesting. But in general classification problems, true negatives often matter a great deal.
Consider a medical screening test. You care whether it correctly identifies sick patients (true positives), and you care whether it correctly identifies healthy patients (true negatives). A test that falsely alarms on every healthy person is a serious problem, but F1 does not capture that failure. This asymmetry is one of the main criticisms in the research literature: because F1 stresses one class, it seems inappropriate for problems where both classes matter equally.3ACM Computing Surveys. A Review of the F-Measure: Its History, Properties, Criticism, and Alternatives
The Matthews correlation coefficient, or MCC, addresses this by incorporating all four quadrants of the confusion matrix: true positives, true negatives, false positives, and false negatives. It produces a high score only when the model performs well across all four categories, scaled proportionally to the sizes of the positive and negative classes. Research has argued that MCC provides a more reliable and informative evaluation of binary classifiers than either F1 or accuracy.4PubMed Central. The advantages of the Matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation Despite this, F1 remains far more commonly reported, partly out of inertia and partly because many practitioners work in domains where the positive class genuinely is the only one that matters.
Prevalence Sensitivity and Cross-Dataset Comparisons
Another subtle issue is that the F1 score depends on the prevalence of the positive class in the dataset. If you train and evaluate a fraud detection model on a dataset where 2% of transactions are fraudulent, and someone else evaluates their model on a dataset where 10% are fraudulent, the F1 scores are not directly comparable. The underlying difficulty of the task changes with prevalence, and F1 reflects that difficulty without making it visible.5Mathematics. Evaluation of the F1* Score Across Twelve Text Datasets with Prevalence Sensitivity Analysis
This is not just a theoretical concern. In practice, researchers often compare F1 scores across papers, datasets, and even tasks as if they were on a common scale. They are not. A model achieving F1 of 0.85 on a task with balanced classes is doing something very different from a model achieving F1 of 0.85 on a task where the positive class is 3% of the data. The second scenario is usually much harder, and the same F1 number can mask that reality. If you are comparing models across different datasets, metrics like AUC (area under the receiver operating characteristic curve) are often more stable, because AUC evaluates model performance across all possible thresholds rather than at a single operating point.
How F1 and AUC Tell Different Stories
F1 and AUC frequently disagree about which model is best, and understanding why helps clarify what each one measures. AUC evaluates how well a model ranks positive examples above negative ones across every possible decision threshold. F1 evaluates how well a model performs at one specific threshold, the point where it actually decides “positive” or “negative.” A model can have excellent AUC (meaning its scores do a good job of separating the classes) but poor F1 (meaning the threshold it happens to be using converts those scores into bad decisions).
A benchmarking study of five machine-learning algorithms on enterprise risk-prediction data with roughly 8.6% adverse-event prevalence illustrates this gap. The top-performing gradient-boosted model achieved an AUC of 0.872 but an F1 of only 0.532. The logistic regression model scored an AUC of 0.781 and an F1 of 0.402.6Review of Applied Science and Technology. Quantitative Benchmarking of Machine Learning Models for Risk Prediction: A Comparative Study Using AUC/F1 Metrics and Robustness Testing The AUC numbers look respectable, but the F1 numbers reveal that at the chosen threshold, neither model is doing a great job of balancing precision and recall for the minority class. The disconnect is not a flaw in either metric; it reflects the fact that AUC and F1 answer fundamentally different questions. AUC asks “how well can this model discriminate?” while F1 asks “how well is this model actually performing right now, at this threshold?”
Choosing the Right Threshold
Because F1 depends on the decision threshold, picking the right threshold matters enormously. Most classifiers output a continuous score or probability rather than a hard yes-or-no label. Somewhere you have to draw a line and say “everything above this score counts as positive.” The default in many software packages is 0.5, but that default is often far from optimal for F1.
There is a clean mathematical result here: if a classifier’s output scores are well-calibrated probabilities, the threshold that maximizes F1 is half the optimal F1 value itself.7PubMed Central. Optimal Thresholding of Classifiers to Maximize F1 Measure So if the best achievable F1 is 0.80, the optimal threshold would be around 0.40, not the 0.50 default. In practice, this means the optimal threshold is almost always lower than 0.5 for problems with rare positive classes. Moving the threshold down means flagging more cases as positive, which boosts recall at the expense of some precision, but on imbalanced datasets this trade-off usually improves F1.
That said, threshold tuning is not always a free lunch. A large-scale study across 45 classification tasks found that on some well-known benchmarks, tuning the threshold gave essentially zero improvement. On a widely-used credit-card fraud dataset, a standard random forest at the default 0.5 threshold already achieved an F1 of about 0.86, and threshold tuning actually decreased F1 by a tiny amount.8arXiv. When Single-Dataset Conclusions Fail: A 45-Task Study of Threshold Tuning and Resampling for Imbalanced Classification The lesson is that threshold optimization helps in many cases but should be validated on held-out data rather than assumed to be universally beneficial.
Training Directly on F1
A practical headache with the F1 score is that it is not smooth. It is computed from discrete counts of true positives, false positives, and false negatives, which means you cannot just take its derivative and plug it into the gradient-based optimization that neural networks rely on. Most deep learning models are trained by minimizing a smooth loss function like cross-entropy, and then F1 is calculated after the fact as an evaluation metric. This creates a disconnect: you optimize for one thing during training but judge the model on something else.
Researchers have worked to bridge this gap by designing differentiable approximations that can serve as stand-in loss functions during training. One approach, called sigmoidF1, replaces the hard counts in the F1 formula with smooth sigmoid functions, creating a loss that neural networks can optimize using standard gradient descent while approximating F1 directly.9arXiv. sigmoidF1: A Smooth F1 Score Surrogate Loss for Multilabel Classification Another line of work uses soft sets and differentiable approximations of the step function to convert confusion-matrix-based metrics, including F1, into trainable objectives.10NeurIPS Proceedings. Differentiable approximation for training neural network binary classifiers These methods have shown promise, particularly in multilabel settings where each example can belong to multiple categories simultaneously, but they remain less commonly used than standard cross-entropy loss followed by post-hoc F1 evaluation.
Getting Confidence Intervals Around F1
A single F1 number tells you almost nothing about how reliable that number is. If you shuffle the test set slightly, would F1 be 0.82 or 0.78? The spread matters, especially when you are comparing two models and trying to decide whether one is genuinely better. Yet many published papers report F1 without any measure of uncertainty.
The most common practical approach is bootstrapping: you resample the test set many times with replacement, compute F1 on each resample, and look at the spread of the resulting scores to form a confidence interval. Some research groups have adopted this practice, computing evaluation metrics across multiple bootstrap samples and running statistical tests to determine whether differences between models are meaningful or just noise.11Natural Language Processing Journal. MISTRA: Misogyny Detection through Text–Image Fusion and Representation Analysis Formal confidence intervals for both micro-averaged and macro-averaged F1 have been developed to put this on more rigorous footing.12SpringerLink (Appl Intell). Confidence interval for micro-averaged F1 and macro-averaged F1 scores If you are reading a paper that claims one model beats another by a point or two of F1 but provides no confidence intervals, treat the claim with skepticism. Small differences can easily fall within the noise.
The Weak Theoretical Foundation
For all its popularity, the F1 score has a surprisingly thin theoretical justification. A comprehensive review of the measure’s history and properties concluded that the rationale underlying F1 appears weak: unlike certain other metrics, it does not have a clear representational meaning, and the choice of the harmonic mean over other possible aggregation functions has little theoretical grounding.13ACM Computing Surveys. A Review of the F-Measure: Its History, Properties, Criticism, and Alternatives The harmonic mean was not selected because it was derived from some principled loss function or decision-theoretic framework. It was selected because it penalizes imbalance, which seemed like a desirable property, and the community adopted it.
This does not mean F1 is useless. In the many settings where you genuinely care only about precision and recall for a specific positive class, F1 gives you a reasonable single-number summary. But when the theoretical underpinning is “it seemed like a good idea and everyone uses it,” you should treat the score as a practical convenience rather than a fundamental truth about your model’s quality. It is one lens, not the lens.
Fairness-Aware Extensions
As machine learning systems are increasingly deployed in high-stakes settings like hiring, lending, and criminal justice, researchers have begun asking whether standard metrics like F1 are sufficient for evaluating models that must be both accurate and fair. A model can achieve high F1 overall while performing much worse on certain demographic groups, a gap that traditional F1 does not surface.
Recent work has proposed fairness-adapted versions of confusion-matrix metrics, including fair-precision, fair-recall, and fair-F1, designed to quantify the trade-off between accuracy and individual fairness. One approach using a Siamese network architecture reported achieving meaningfully higher individual fairness alongside higher accuracy compared to existing bias-mitigation methods, with fair-F1 improvements of roughly five to ten percentage points.14AAAI Publications. Accurate Fairness: Improving Individual Fairness without Trading Accuracy These metrics are still in early stages of adoption, but they point toward a future where evaluation is not just about overall performance but about how evenly that performance is distributed across the people the model affects.
F1 for Measuring Human Agreement
An interesting use of F1 that extends beyond model evaluation is measuring how much human annotators agree with each other. When you build a labeled dataset for machine learning, multiple people typically label the same examples, and you need to know how consistent they are. Traditional agreement metrics like Cohen’s kappa or Krippendorff’s alpha have limitations, particularly when annotators label only a subset of the data or when the task involves multiple labels per example.
Researchers have proposed adapting the F1 score as an agreement metric, treating one annotator’s labels as the “predictions” and another’s as the “ground truth” and computing F1 between them. This approach supports binary, multi-class, and multi-label data, accommodates incomplete annotations, and has the appealing property that you can directly compare human-to-human agreement with model-to-human agreement on the same scale.15Frontiers in Artificial Intelligence and Applications. Underperformance or Pluralism: A Machine Learning Perspective on Inter-Annotator Agreement If two human annotators achieve an F1 of 0.80 against each other on a labeling task, a model achieving F1 of 0.78 against the gold-standard labels is performing close to the human ceiling. Without that context, the 0.78 might look disappointing when it is actually about as good as the task allows.

