Uppsats

Machine Learning for Predictive Maintenance in Real Estate : Forecasting Annual Maintenance Events and Costs using Historical Transaction Logs

Yrkesexamen på avancerad nivå

Umeå universitet/Institutionen för matematik och matematisk statistik

Publicerad: 2026

Språk: Engelska

Sammanfattning

Traditional real estate maintenance budgeting heavily relies on static, calendar-based schedules, which often results in operational inefficiencies and suboptimal financial allocations. In this thesis, the application of machine learning is investigated to transition from this reactive approach toward a data-driven predictive maintenance strategy. The primary objective is to forecast annual component-level maintenance events and their associated costs using historical transaction logs and property characteristics, ultimately generating reliable budget predictions that can outperform manual budgeting practices. The thesis utilized a dataset comprising over 100,000 historical maintenance records from a Swedish residential property portfolio, integrated with structural property attributes and adjusted for inflation. Due to high zero-inflation in the data, where 73.2% of observations recorded zero maintenance costs, a Two-Stage Hurdle model was implemented. The first stage utilized classification models, including Logistic Regression, Random Forest, and CatBoost, to predict the probability of a maintenance event. The second stage applied regression algorithms, specifically Ordinary Least Squares, Random Forest, and CatBoost, to estimate the financial cost per square meter. Model performance was evaluated using an expanding-window time-series cross-validation approach. Model selection in the classification stage was based primarily on the F1-score and ROC AUC, since both are robust under the strong class imbalance in the data, while the percentage deviation between the predicted and actual portfolio cost assessed the budgeting stage. Under these criteria, the three classifiers performed at a broadly comparable level, with F1-scores between 0.700 and 0.704 and differences that were not statistically significant. The clearest distinction emerged at the portfolio-budgeting stage: Random Forest produced the lowest property-level errors but underestimated the total portfolio cost by 18.63%, whereas CatBoost achieved a near-neutral portfolio deviation of -2.65%, an 8.9% improvement over the OLS baseline. These findings indicate that machine learning models trained on historical transaction data can serve as a viable alternative to manual maintenance budgeting, while the performance gaps between the individual algorithms remain modest.

Information

Lärosäte / institution
Umeå universitet/Institutionen för matematik och matematisk statistik
Publiceringsdatum
2026
Uppsatstyp
Yrkesexamen på avancerad nivå
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.