Customer Personality Data for Predicting Consumer Response in Digital Business
Keywords:
Digital Business Campaigns, Bald Eagle Search, Random Forest, Feature Importance Analysis, Consumer Response PredictionAbstract
The rapid growth of digital business has transformed customer engagement, shifting from traditional marketing approaches toward highly personalized, data-driven strategies. Understanding consumer behavior is therefore essential for sustaining competitiveness. This study investigates the use of customer personality data to predict consumer responses in digital business campaigns by comparing four machine learning scenarios: Bald Eagle Search–Random Forest (BES-RF), Bald Eagle Search–KNN (BES-KNN), baseline Random Forest (RF), and baseline KNN. The dataset includes demographic attributes, purchasing behavior, and historical campaign acceptance, offering a comprehensive view of customer characteristics. Evaluation results show that BES-RF achieved the highest accuracy (88.1%) and AUC (0.86), with relatively strong precision (70.3%) but limited recall (27.4%), reflecting challenges in capturing minority classes. RF performed similarly (accuracy 87.6%, AUC 0.85), while BES-KNN improved precision (76.2%) but suffered from very low recall (16.8%). KNN recorded the weakest balance across metrics (accuracy 86.5%, recall 22.1%). Macro-averaged metrics highlight poor sensitivity to responders, weighted averages inflate performance due to majority dominance, and Matthews Correlation Coefficient values confirm only moderate predictive consistency. Confidence intervals for accuracy and AUC suggest that improvements over baselines are modest and not definitive. Consumer pattern analysis revealed that higher income, recent purchases, and product-specific expenditures particularly wine and meat products were influential predictors of responsiveness. However, the current evidence does not confirm effectiveness. The BES algorithm is insufficiently defined, baseline comparisons are limited, validation is unclear, and recall remains very low. Limitations include reliance on a single secondary dataset, class imbalance, possible leakage and outliers, cross-sectional design, lack of external validation, limited generalizability, and inability to infer causality. Future work should focus on methodological improvements such as resampling techniques, cost-sensitive learning, or hybrid ensembles to improve recall without sacrificing precision. Expanding datasets to include psychographic attributes, social media engagement, and real-time behavioral signals, alongside external validation and longitudinal designs, would enrich predictive accuracy and strengthen generalizability.




