A random forest classifier is a machine-learning algorithm that makes predictions by building a large collection of decision trees and combining their individual votes into a single answer. Instead of relying on one tree that might overfit quirks in the training data, the forest pools hundreds or thousands of independently grown trees, each trained on a slightly different slice of the data and a random subset of the available features. The result is an algorithm that tends to be accurate, resistant to noise, and relatively forgiving when you haven’t perfectly tuned every setting. That combination has made random forests one of the most widely used classifiers across fields from genomics to land-cover mapping, though the method has real blind spots worth understanding before you trust it with a high-stakes decision.
How the Forest Gets Built
Each tree in a random forest starts from a bootstrapped sample of the training data, meaning the algorithm draws a random sample (with replacement) that is the same size as the original dataset. Some observations appear more than once in a given sample, and roughly a third get left out entirely. That left-out portion turns out to be useful later for estimating how well the model is performing.
At every node of a growing tree, the algorithm considers only a random subset of the available features rather than all of them. It picks whichever feature and split point best separates the classes at that node, then moves on. By randomizing both the data each tree sees and the features it can choose from, the forest ensures that individual trees are different from one another. Trees that disagree on easy cases will still tend to agree on clear patterns, while their idiosyncratic errors cancel out when votes are tallied. For a classification task the final prediction is simply the class that receives the most votes across all trees.
How Each Split Is Chosen
Inside every tree, the algorithm needs a rule for deciding which feature to split on and where. The two most common criteria are the Gini index and entropy, both of which measure how “mixed” the classes are at a given node. A split that sends most of one class to the left branch and most of the other class to the right branch scores well on either metric. Traditional tree-growing algorithms like CART and C4.5 rely on these impurity-reduction functions to promote the discriminative power of each split, and research has shown that simple enhancements to those functions can also provide guarantees about how compact the resulting trees will be.
1PubMed Central. Regularized impurity reduction: accurate decision trees with complexity guaranteesIn practice, the choice between Gini and entropy rarely makes a dramatic difference to the forest’s overall accuracy. What matters more is the number of features randomly sampled at each split and the total number of trees. Increasing the number of trees almost always improves performance up to a point of diminishing returns, while the feature-subset size controls how correlated the individual trees are with each other. A smaller subset means more randomness, less correlation, and often better generalization, though too small a subset starves trees of useful information.
Measuring Error Without Holding Data Back
One of the most practical conveniences of a random forest is its built-in error estimate. Because each tree is trained on a bootstrap sample, every observation gets left out of some fraction of the trees. The algorithm can run each observation through only the trees that never saw it during training, tally their votes, and check whether the prediction is correct. The resulting “out-of-bag” (OOB) error gives you a performance estimate without ever needing a separate validation set.
The OOB error is often described as an unbiased estimate of the true error rate, and it is computationally cheap because you only need to build one forest rather than running repeated cross-validation. That said, the picture is not quite that clean. For two-class problems, the OOB error can actually overestimate the true prediction error depending on parameter choices, meaning your model may be slightly better than the OOB number suggests.
2PubMed Central. On the overestimation of random forest’s out-of-bag errorThe practical takeaway is that OOB error is a solid quick-and-dirty gauge, especially when you’re iterating fast and don’t want to spend time on full cross-validation. But for a final reported accuracy number, especially in a publication or a deployment decision, running proper cross-validation on top of the OOB estimate is worthwhile.
Variable Importance and Its Pitfalls
Beyond making predictions, random forests offer a way to rank which features matter most. The two standard approaches are permutation importance, which measures how much accuracy drops when a feature’s values are randomly shuffled, and impurity-based importance, which adds up how much each feature contributes to reducing the Gini index or entropy across all splits in all trees. Both are useful, but both can mislead you.
Simulation studies have demonstrated that when the data contains variables of different types, such as categorical features with many levels alongside continuous features, the standard importance measures can be biased. Variables with more categories or a wider range of values get artificially preferred, not because they are truly more predictive but because the splitting mechanism has more ways to exploit them. Two mechanisms drive this bias: biased variable selection within individual trees and effects induced by the bootstrap sampling itself.
3PubMed Central. Bias in random forest variable importance measures: illustrations, sources and a solutionThe Gini-based importance measure is particularly susceptible. When features vary in their minor allele frequency or their number of categories, Gini importance tends to increase with the number of possible split points, even for features with no real predictive power. Corrected versions of the importance measure, such as the “Actual Impurity Reduction” approach, address this by adjusting for the bias, producing values that center around zero for irrelevant features regardless of their type.
4Bioinformatics. The revival of the Gini importance?If you’re using variable importance to decide which features to investigate further, whether that means running lab experiments or collecting more data, this bias can send you chasing the wrong leads. Using a corrected importance measure or a permutation-based approach with a held-out test set is a simple safeguard.
Going Deeper With SHAP
Standard variable importance tells you which features matter globally, across all predictions, but it doesn’t explain why a particular observation received its specific prediction. That’s where SHAP (SHapley Additive exPlanations) values come in. SHAP assigns each feature a contribution score for every individual prediction, letting you see not just that, say, blood glucose was important on average, but that this patient’s unusually high glucose level pushed their risk score up by a specific amount.
Applied to random forest models trained on metabolomics data, tree-based SHAP has been shown to provide more explanatory depth than traditional approaches like VIP scores from partial-least-squares models. SHAP can show directionality, showing whether a high value of a feature increases or decreases the predicted class probability, which standard importance measures do not reveal.
5PubMed Central. Interpretable machine learning with tree-based shapley additive explanations: Application to metabolomics datasets for binary classificationFor anyone building a random forest that needs to be explained to stakeholders, whether that’s a clinical team reviewing a diagnostic model or a regulatory agency evaluating a credit-scoring system, SHAP has become the de facto standard for post-hoc interpretation.
Dealing With Imbalanced Classes
Random forests can struggle when one class vastly outnumbers the other in the training data. If 95% of your samples belong to the majority class, a forest can achieve 95% accuracy simply by predicting the majority class for everyone, which is useless if you’re trying to catch the rare 5%. The model doesn’t deliberately ignore the minority class; it just doesn’t see enough examples to learn their patterns well.
Several strategies address this. One common approach uses repeated random sub-sampling to create balanced training sets: you split the majority class into multiple subsets, pair each subset with the full minority class, and train a separate forest on each balanced pair. The final prediction aggregates across all of these forests.
6PubMed Central. Predicting disease risks from highly imbalanced data using random forestA related technique, the Balanced Random Forest, adjusts the bootstrap step so that each tree’s training sample is balanced. Modifications to this idea apply clustering-based under-sampling at each bootstrap iteration, further refining which majority-class observations end up in each tree.
7International Journal of Advances in Intelligent Informatics. Modified balanced random forest for improving imbalanced data predictionWhich method works best depends on how severe the imbalance is and how much data you have overall. For moderate imbalances, simply adjusting class weights within the forest’s splitting criterion is often enough. For extreme imbalances, such as predicting rare diseases from electronic health records, the sub-sampling and balanced-forest approaches are worth the extra complexity.
Handling Missing Values
Real-world datasets are messy, and missing values are the norm rather than the exception. Random forests have a somewhat unusual advantage here: several algorithms exist that let the forest handle missing data internally, without requiring you to fill in blanks beforehand. Approaches include proximity-based imputation, where the forest uses the similarity structure among observations to fill in gaps, and “on-the-fly” imputation, which handles missing values during the tree-growing process itself. A newer approach called missForest uses multivariate supervised splitting to impute missing values and has become popular as a general-purpose imputation tool, even for data that will ultimately be analyzed by a completely different method.
8PubMed Central. Random Forest Missing Data AlgorithmsThe practical advantage is that you can often throw messy data at a random forest without an elaborate preprocessing pipeline. That said, these internal imputation methods do add computation time and can introduce their own biases, so for high-stakes applications it’s still worth understanding which values are missing and why before relying on the forest to sort it out.
Where Random Forests Excel in Practice
Random forests have found a home in a remarkably wide range of fields. In genomics, where a typical microarray experiment might measure the activity of tens of thousands of genes across a handful of patient samples, random forests perform comparably to support vector machines and other classifiers while yielding very compact sets of informative genes. Studies on microarray data have shown that random-forest-based gene selection can produce smaller feature sets than alternative methods while maintaining predictive accuracy.
9PubMed Central. Gene selection and classification of microarray data using random forestIn remote sensing, random forests are a workhorse for turning satellite imagery into land-cover maps. An assessment of the method’s effectiveness for land-cover classification found that random forests achieved about 92% overall accuracy and significantly outperformed single decision trees. The model also proved robust to noisy data and reduced training sets: significant drops in performance appeared only when training data was cut by more than half or noise exceeded 20%.
10ISPRS Journal of Photogrammetry and Remote Sensing. An assessment of the effectiveness of a random forest classifier for land-cover classificationOn specific feature-selection tasks in bioinformatics, combining random forest importance scores with forward- and backward-search strategies has achieved classification accuracies as high as 100% on leukemia datasets and in the range of 85% to 96% on breast, prostate, nervous-system, and lymphoma datasets.
11PubMed. A novel random forests-based feature selection method for microarray expression data analysisBeyond standard classification, the random forest framework has been extended to handle survival data, where the outcome is the time until an event like death or relapse. Random survival forests adapt the splitting and aggregation rules to accommodate censored observations, providing a nonparametric alternative to classical survival models.
12The Annals of Applied Statistics. Random survival forestsScalability and Speed
Random forests scale reasonably well, but they aren’t free. Each tree needs to sort and split the data, and you’re building hundreds or thousands of them. The time complexity for building a forest of m trees over N samples is roughly proportional to m × N × log(N), which means doubling your dataset doesn’t quite double your training time, but adding more trees always adds more wall-clock minutes.
13ScienceDirect. Exploring time complexity and machine learning scalability for COVID-19 Predictions: A case study from Saudi ArabiaFor moderately sized datasets, say up to a few hundred thousand rows and a few hundred features, a random forest typically trains in minutes to hours on a standard machine. Once you’re dealing with millions of rows or very high-dimensional data, training times can stretch to days. The trees are inherently parallelizable, though, since each one is independent, so throwing more CPU cores at the problem offers nearly linear speedups.
Prediction time is also worth considering. At inference, every new observation has to traverse every tree. A forest of a thousand trees, each a hundred nodes deep, means a hundred thousand node-traversals per prediction. For batch predictions this is fine, but for latency-sensitive applications like real-time fraud detection, it can be a bottleneck.
When Random Forests Fail to Generalize
A common misconception is that random forests are universally robust. They are robust to many things, including noise, moderate overfitting, and correlated features, but they have a fundamental weakness when it comes to extrapolation. A random forest cannot predict values outside the range of what it has seen in training. If you train a forest on temperatures between 10°C and 30°C, it cannot predict what will happen at 40°C because every leaf node’s vote is bounded by the training data range.
This limitation becomes especially acute in spatial applications. A study on environmental modeling found that random forests with spatial proxies were poorly suited for spatial extrapolation to new areas. When predicting in regions that differed geographically from the training domain, baseline models consistently outperformed the spatial-proxy models because a very large part, sometimes all, of the extrapolation area had feature values not covered by the training data.
14Geoscientific Model Development. Random forests with spatial proxies for environmental modelling: opportunities and pitfallsIf your problem involves predicting into genuinely new territory, whether geographic, temporal, or in feature space, a random forest will silently return its best guess from the nearest known region without any warning that it’s operating outside its comfort zone. Pairing the forest with uncertainty estimates or using models designed for extrapolation is the safer path in those situations.
Shrinking Forests for Small Devices
Deploying a random forest on a server is straightforward, but putting one on an embedded device, a smartphone’s edge processor, or an IoT sensor is a different story. A forest of 500 trees, each with thousands of nodes, can consume megabytes of memory that a microcontroller simply doesn’t have. Prediction speed also matters when battery life is limited and latency requirements are tight.
Recent work on model compression addresses this directly. One approach, called Splitting Stump Forests, reduces model size and inference time by selecting nodes that evenly partition the incoming data and then fitting a lightweight linear model on the resulting representation. Empirical evaluations show that these compressed models outperform full random forests and other state-of-the-art compression methods when deployed on memory-limited embedded devices.
15Machine Learning. Splitting stump forests: tree ensemble compression for edge devices (extended version)This is a growing area of research because the use cases are compelling: wildlife monitoring sensors in remote locations, wearable health devices, industrial equipment running anomaly detection. In each case, you want the accuracy of an ensemble model without the memory and power footprint of the full forest.
Random Forests Compared to Other Approaches
In the broader landscape of machine learning, random forests occupy a particular niche. Compared to a single decision tree, a random forest is almost always more accurate. The land-cover study mentioned earlier confirmed this with a formal statistical test, finding significantly better performance from the forest at an extreme significance level.
16ISPRS Journal of Photogrammetry and Remote Sensing. An assessment of the effectiveness of a random forest classifier for land-cover classificationCompared to gradient-boosted trees (like XGBoost or LightGBM), random forests tend to be easier to tune and more resistant to overfitting but often slightly less accurate on structured tabular data when the boosted model is well optimized. The trees in a random forest are grown independently and in parallel, while boosted trees are grown sequentially, with each new tree correcting the errors of the ones before it. That sequential dependence gives boosting more power to fit complex patterns but also more rope to overfit.
Compared to deep neural networks, random forests have the advantage of requiring much less data and much less computational infrastructure. They work well on tabular data out of the box, while neural networks typically need careful architecture design and extensive tuning to match forest performance on non-image, non-text tasks. On the other hand, neural networks excel at learning from images, audio, and text in ways that random forests simply cannot.
The situations where random forests shine brightest tend to share a few characteristics: tabular or structured data, moderate dataset sizes, a mix of feature types, a need for interpretable importance rankings, and limited time or resources for hyperparameter tuning. If your problem fits that profile, a random forest is often the strongest starting point, even if you ultimately move to a more complex model later.

