How Does Random Forest Regression Work?

Random forest regression works by growing hundreds of decision trees on slightly different slices of your data, then averaging all their individual predictions into a single, more stable output. Each tree sees a random subset of the training observations and, at every split, considers only a random handful of the available input features. That deliberate injection of randomness is what makes the combined prediction far more reliable than any single tree could manage on its own, and it is the reason random forests have become one of the most widely used tools for predicting continuous outcomes across fields from ecology to drug design.

How the Algorithm Builds Its Predictions

A random forest for regression starts by drawing many bootstrap samples from your training data. A bootstrap sample is the same size as the original dataset but drawn with replacement, so some observations appear more than once while others are left out entirely. For each bootstrap sample, the algorithm grows a full decision tree. At each node of that tree, instead of evaluating every available feature to find the best split, the algorithm randomly selects a small subset of features and picks the best split from that limited menu. Once all the trees are built, a new data point is fed through every tree, and the final prediction is simply the average of all the individual tree outputs.1Geoderma. Comparing the prediction performance, uncertainty quantification and extrapolation potential of regression kriging and random forest while accounting for soil measurement errors

This two-layer randomness, randomizing both which observations each tree trains on and which features each split can use, is the heart of the algorithm. Each individual tree is a mediocre predictor. It has seen only part of the data and has been forced to ignore most of the features at any given decision point. But the errors of all those mediocre trees tend to point in different directions, so when you average them, the mistakes cancel out and what remains is a strong, stable prediction.

Why Adding Randomness Improves Accuracy

It sounds counterintuitive that deliberately handicapping each tree would make the ensemble better. The key insight is that randomness reduces the correlation between trees. If every tree in the forest were allowed to pick the globally best feature at every split, most trees would end up looking nearly identical, especially when a few features dominate the signal. Averaging a hundred nearly identical predictions does little to reduce error. By forcing each tree to work with a random subset of features, the algorithm ensures that different trees discover different patterns in the data, and averaging genuinely diverse predictions produces a much tighter estimate.

Research has shown that this random feature selection does more than just reduce the variability of predictions. It can also reduce systematic bias. When the signal-to-noise ratio in the data is high, random forests outperform plain bagging ensembles (which randomize observations but not features) because the feature randomization helps the forest capture what researchers call “hidden patterns,” relationships in the data that bagged ensembles struggle to find.2arXiv. Randomization Can Reduce Both Bias and Variance: A Case Study in Random Forests In other words, the randomness is not just a variance-reduction trick. It genuinely helps the model learn structure it would otherwise miss.

One recent paper drew a striking analogy: the way genetically identical ants in a colony achieve functional differentiation through stochastic responses to local cues maps precisely onto the bootstrap aggregation and random feature subsampling that decorrelate decision trees in a random forest.3arXiv. Decorrelation, Diversity, and Emergent Intelligence: The Isomorphism Between Social Insect Colonies and Ensemble Machine Learning Each ant (or each tree) is a simple, noisy decision-maker, but the colony (or the forest) exhibits robust, intelligent behavior because the individual agents are sufficiently different from one another.

Interpreting Which Features Matter

One of the most common reasons people reach for random forests is the built-in variable importance measure. After the forest is trained, you can ask: which features contributed most to prediction accuracy? The standard approach is permutation importance. You randomly shuffle the values of a single feature across the dataset and measure how much the prediction error increases. A feature that, when scrambled, causes a large jump in error was clearly doing a lot of work.

This is genuinely useful, but it has a well-documented blind spot with correlated features. When two input variables are strongly correlated, the standard permutation importance measure tends to inflate the importance of both. Two separate mechanisms drive this bias: the tree-building process itself prefers correlated predictors at the splitting stage, and the unconditional permutation scheme gives correlated variables an additional artificial boost.4PubMed Central. Conditional variable importance for random forests When predictors are correlated and genuinely associated with the outcome, the standard importance measure overstates the contribution of the correlated group. Conditional variable importance measures, which account for the correlation structure, produce less distorted rankings.5PubMed Central. The behaviour of random forest permutation-based variable importance measures under predictor correlation

In practice, many analysts now supplement or replace permutation importance with SHAP values, a model-agnostic interpretability tool. SHAP assigns each feature a contribution to each individual prediction, rather than giving a single global ranking. Combined with partial dependence plots, this approach can reveal nonlinear, locally varying relationships that a global importance ranking would miss. A groundwater contamination study in China, for instance, used SHAP alongside partial dependence plots to show that the relationship between groundwater nitrate and predictors like soil moisture and groundwater depth was far more complex than a simple positive-or-negative association.6Physics and Chemistry of the Earth, Parts A/B/C. Predicting groundwater nitrate contamination and identifying key drivers in southern Weinan, northwest China using random forest and SHapley additive exPlanations (SHAP) approach If you are using random forest regression to understand what drives your outcome, not just to predict it, investing in these richer interpretation tools is worth the effort.

The Extrapolation Problem

Every decision tree makes predictions by routing a data point to a leaf node and returning the average of the training observations that landed in that leaf. This means a tree can never predict a value higher than the highest training value it saw, or lower than the lowest. A random forest, being an average of trees, inherits the same constraint. If your training data contains house prices ranging from $100,000 to $900,000 and you ask the model to predict a house that would sell for $1.2 million, the forest will cap its prediction somewhere near the top of its training range.

This inability to extrapolate beyond the training distribution is one of the most important practical limitations of random forest regression. It rarely matters when you are interpolating, predicting within the range and feature space the model has already seen. But in fields like drug design, where the whole point is to find compounds with higher activity than anything tested so far, standard random forest regression can be fundamentally limiting. Research comparing standard regression against a pairwise reformulation across thousands of drug design and materials property datasets found that the pairwise approach, which reframes the problem as predicting differences between compounds rather than absolute activity, vastly outperformed standard regression at extrapolation tasks.7Machine Learning. Extrapolation is not the same as interpolation

If your use case requires predicting values outside the range of historical data, you should be aware that random forests will quietly clip those predictions. Gradient boosting models share this limitation to a lesser degree (they also build trees, after all), and linear models, despite their simplicity, at least produce predictions that extend the learned trend beyond the training range. The right tool depends on whether your problem is fundamentally about interpolation or extrapolation.

Evaluating Accuracy with Out-of-Bag Estimates

Because each tree is trained on a bootstrap sample, roughly a third of the training observations are left out of any given tree. These leftover observations are called out-of-bag (OOB) samples. You can run each observation through only the trees that did not use it during training, average those predictions, and compare them to the actual values. The result is an error estimate that comes essentially for free, without needing a separate holdout set or cross-validation.

This convenience has made OOB error the default evaluation metric in many random forest implementations, but it deserves some caution. For classification problems with continuous input features, the OOB error can overestimate the true prediction error depending on the tuning parameter settings. Stratified subsampling with sampling fractions proportional to the class sizes produces less biased error estimates.8PubMed Central. On the overestimation of random forest’s out-of-bag error For regression tasks specifically, the OOB residuals (the differences between actual values and OOB predictions) can be used to consistently estimate the residual variance, giving you a handle on how much irreducible noise remains after the model has done its best.9Statistics & Probability Letters. Consistent estimation of residual variance with random forest Out-Of-Bag errors

The practical takeaway: OOB error is a good quick diagnostic, but for any final model evaluation, especially when tuning hyperparameters, confirm the results with an independent test set or proper cross-validation. Relying solely on OOB metrics can lead you to choose slightly suboptimal settings.

How Random Forests Compare to Gradient Boosting

The most common alternative to random forest regression is gradient boosting, which also builds an ensemble of trees but does so sequentially. Each new tree in a gradient boosting model is trained to correct the errors of all the trees that came before it. This sequential, error-correcting strategy often gives gradient boosting a slight edge in raw accuracy on tabular data, which is why algorithms like XGBoost and LightGBM dominate machine learning competitions.

In practice, the accuracy gap is often smaller than people expect. A study predicting tree canopy cover across four pilot regions found that random forests and stochastic gradient boosting produced nearly identical test-set errors, with mean squared errors identical to three decimal places in all four regions. Where the two methods differed more was in interpretability and ease of use: gradient boosting tended to concentrate variable importance in fewer features, while random forests spread importance more evenly. Random forests were also simpler to implement because they have fewer parameters to tune and are less sensitive to the exact values chosen for those parameters.10Canadian Journal of Forest Research. Random forests and stochastic gradient boosting for predicting tree canopy cover: comparing tuning processes and model performance

In genomic selection, where the task is to predict breeding values from genetic markers, boosting outperformed both random forests and support vector machines, with correlations between predicted and true breeding values of 0.547 for boosting compared to 0.483 for random forests.11PubMed Central. A comparison of random forests, boosting and support vector machines for genomic selection Results like these suggest that when squeezing out the last percentage point of accuracy matters and you have the time and expertise to tune a more complex model, gradient boosting is often the better choice. But when you need a robust baseline that works well out of the box, random forests remain hard to beat.

Tuning the Knobs

Random forests have several settings that you, the user, need to choose. The most important ones are:

  • Number of trees: More trees generally improve stability but increase computation time. Past a few hundred trees, accuracy gains typically flatten out.
  • Features per split: The number of randomly selected features considered at each split. For regression, the common default is one-third of the total features. Lower values increase diversity between trees; higher values let each tree be individually stronger.
  • Minimum node size: How many observations must remain in a leaf for the tree to stop splitting. Smaller leaves mean more flexible trees; larger leaves act as regularization.
  • Sampling strategy: Whether bootstrap samples are drawn with or without replacement, and what fraction of the data each tree sees.

The good news is that random forests are famously tolerant of their default hyperparameter values. In most cases, the algorithm works reasonably well straight out of the box. That said, tuning these parameters, especially the number of features per split, can meaningfully improve performance.12WIREs Data Mining and Knowledge Discovery. Hyperparameters and tuning strategies for random forest Grid search or random search over a modest range of values is usually sufficient. Unlike gradient boosting, where the learning rate, tree depth, and number of iterations interact in complex ways, random forest tuning is forgiving. A slightly off setting for one parameter rarely tanks the model.

Handling Missing Data

Real-world datasets are messy, and missing values are the norm rather than the exception. Random forests have a natural advantage here. Because each split in a tree only looks at one feature at a time, there are straightforward ways to route an observation with a missing value: send it down both branches and weight the results, or use a surrogate split based on a correlated feature. Several algorithms go further and use random forests themselves as the imputation engine.

The family of random forest missing data algorithms includes proximity-based imputation, on-the-fly imputation during tree construction, and multivariate approaches that use both supervised and unsupervised splitting to fill in gaps. These methods handle mixed data types (continuous and categorical variables side by side), adapt naturally to nonlinear relationships and interactions, and scale to large datasets.13PubMed Central. Random Forest Missing Data Algorithms One algorithm in this family, missForest, has become a popular general-purpose imputation tool even outside the random forest prediction context. You train a random forest to predict each variable with missing values from all the other variables, iterate until the imputations stabilize, and end up with a complete dataset ready for any downstream analysis.

This is a genuinely useful property. Many competing regression methods require you to handle missing data before modeling, either by dropping incomplete cases (which wastes data) or by running a separate imputation step. Random forests can fold the missing data problem into the modeling itself, which reduces the pipeline complexity and sometimes improves results by letting the imputation and prediction learn from each other.

Applications in Spatial and Environmental Modeling

Random forest regression has found an especially natural home in spatial and environmental sciences, where the goal is often to predict a continuous quantity across a landscape using a patchwork of satellite imagery, terrain data, and climate variables. Soil organic carbon mapping is a good example of this. In one study across Central Vietnamese forests, researchers collected 104 soil samples and used 33 environmental covariates derived from satellite imagery, radar, elevation models, and climate data. The random forest model, when combined with kriging of its residuals (an approach called random forest regression kriging), achieved strong predictive accuracy. Adding more types of environmental data consistently improved the model, and the kriging step, which accounts for spatial patterns in the errors the forest leaves behind, produced improvements in accuracy ranging from roughly 8% to 65% depending on the data combination and the metric used.14Modeling Earth Systems and Environment. Random forest regression kriging modeling for soil organic carbon density estimation using multi-source environmental data in central Vietnamese forests

A similar pattern emerged in a study mapping soil organic carbon in northern Nigeria, where integrating both spectral satellite variables and terrain variables into a random forest model produced the best results, with the combined model achieving the highest explanatory power among the configurations tested.15Artificial Intelligence in Geosciences. Spatial mapping and modelling of soil organic carbon using random forest and remote sensing variables in part of Kaduna, Northern Nigeria Terrain attributes like slope curvature turned out to be unexpectedly influential drivers of carbon distribution, a finding that the random forest’s variable importance measures helped surface.

These environmental applications highlight a broader pattern. Random forest regression tends to shine when you have many input variables of mixed types, when the relationships between inputs and the outcome are nonlinear and full of interactions, and when you want to get a solid predictive model up and running without spending weeks on feature engineering and hyperparameter tuning. The algorithm’s built-in variable importance tools then help you figure out which inputs actually matter, turning a black-box prediction into something closer to scientific understanding. That combination of predictive power and interpretability, even if the interpretability requires some care around correlated features, is why random forests have remained a workhorse algorithm long after flashier alternatives have appeared.