Title: Fuzzy forests for feature selection in high-dimensional survey data: an application to the 2020 US presidential election
Authors: Sreemanti Dey; R. Michael Alvarez
Addresses: California Institute of Technology, Pasadena, CA 91125, USA ' California Institute of Technology, Pasadena, CA 91125, USA
Abstract: An increasingly common methodological issue in the field of social science is high-dimensional and highly correlated datasets that are unamenable to the traditional deductive framework of study. Analysis of candidate choice in the 2020 presidential election is one area in which this issue presents itself: in order to test the many theories explaining the outcome of the election, it is necessary to use data such as the 2020 Cooperative Election Study Common Content, with hundreds of highly correlated features. We present the Fuzzy Forests algorithm, a variant of the popular Random Forests ensemble method, as an efficient way to reduce the feature space in such cases with minimal bias, while also maintaining predictive performance on par with common algorithms like Random Forests and logit. Using Fuzzy Forests, we isolate the top correlates of candidate choice and find that partisan polarisation was the strongest factor driving the 2020 presidential election.
Keywords: fuzzy forests; machine learning; ensemble methods; dimensionality reduction; American elections; candidate choice; correlation; partisanship; issue voting; Trump; Biden.
DOI: 10.1504/IJGUC.2026.152667
International Journal of Grid and Utility Computing, 2026 Vol.17 No.2, pp.126 - 135
Received: 16 Aug 2022
Received in revised form: 09 Nov 2022
Accepted: 14 Nov 2022
Published online: 07 Apr 2026 *