Predicting HVAC Failures
with a Neural Network
A feed-forward network that learns whether a commercial rooftop unit will fail within 30 days — watch it think, watch it get one wrong, and watch backpropagation teach it better.
Rooftop units fail without warning. The data that predicts it already exists.
An unexpected RTU failure at a commercial site means emergency labor rates, downtime, and lost product. Every routine service call already records the features that predict failure — age, maintenance recency, heat load, fault history. A neural network can learn the pattern.
1Supervised Learning
The network trains on labeled history: past service calls where the outcome — failed or didn't — is already known. It learns a function from features to outcomes (Goodfellow et al. 2016, ch. 5).
2The Architecture
4 input features → Hidden Layer 1 (4 neurons, ReLU) → Hidden Layer 2 (3 neurons, ReLU) → 2 output classes. Small enough to see every connection; deep enough to learn feature interactions.
3Why Depth Matters
Layer 1 detects simple feature combinations; layer 2 composes them into concepts like "wear & neglect" or "thermal stress" — hierarchical representation learning (LeCun, Bengio & Hinton 2015).
Five labeled service-call records
Illustrative data patterned on real commercial RTU field history. Each row is a training example: a feature vector paired with a ground-truth label.
| Unit | Age (yrs) | Days Since Service | Avg Ambient (°F) | Fault Codes (90d) | Outcome |
|---|---|---|---|---|---|
| RTU-01 | 12 | 210 | 98 | 4 | FAILED |
| RTU-02 | 3 | 45 | 91 | 0 | DID NOT FAIL |
| RTU-03 | 8 | 120 | 103 | 2 | FAILED |
| RTU-04 ★ | 15 | 30 | 95 | 1 | DID NOT FAIL |
| RTU-05 | 6 | 300 | 99 | 3 | FAILED |
★ RTU-04 is the trap: a very old unit that was just serviced. Age screams risk, recency says safe. Click the row — let's push it through the network.
Run the network yourself
Line thickness = connection weight |w|. Fire the forward pass to see the prediction, reveal what actually happened, then apply backpropagation and watch the weights re-learn in real time.
Each neuron computes z = Σ(wᵢ·xᵢ) + b, then applies ReLU: f(z) = max(0, z). ReLU was chosen over sigmoid because it avoids saturating gradients, keeping the error signal strong for the learning step — a deliberate, theory-rooted design decision.
The error is the fuel
The network predicted WILL FAIL (74%). RTU-04 did not fail. That 0.74 error propagates backward through every connection — and each weight learns exactly its share of the blame.
▼ Age → h1.1 · WEAKENED
Gradient attribution assigns this connection the largest blame: it drove a confident wrong answer. Its weight shrinks proportionally to ∂Error/∂w.
▲ Service Recency → h1.2 · STRENGTHENED
Recent service was the true protective signal the network under-weighted. Its gradient pushes the weight up.
▼ "Wear & Neglect" → FAIL · MODERATED
The concept neuron stays useful, but its vote toward failure is tempered so age alone can no longer dominate the verdict.
Backpropagation uses the chain rule to compute ∂Error/∂w for every weight; gradient descent then updates each one: w ← w − η·∂E/∂w. The learning rate η controls step size — too large overshoots, too small crawls. One example nudges the weights; over many examples the pattern converges. After learning, re-running RTU-04 yields P(fail) ≈ 0.41 → the prediction flips to WILL NOT FAIL. The network learned that a freshly-serviced old unit ≠ a neglected old unit — an interaction no single weight could express.
What this simulation demonstrates
Q1Theory → connection strengths
Initial weights encoded domain priors; then weighted-sum theory defined influence, ReLU kept gradients alive, and every strengthen/weaken decision after the miss followed gradient attribution — not intuition.
Q2What the error taught us
A 74%-confident miss exposed over-reliance on one feature. Backprop converted a single error into targeted, proportional updates. Next: tune η, add training data, hold out a validation set against overfitting.
Q3Real-world applications
Predictive maintenance on RTUs & chillers · cleanroom environmental-monitoring anomaly detection (pharma/biotech) · smart-building BAS optimization · live dispatch prioritization from fault-code telemetry.
Theoretical foundations
- Goodfellow, I., Bengio, Y., & Courville, A. (2016). Deep Learning. MIT Press. deeplearningbook.org — feedforward networks (ch. 6), training (ch. 8), supervised learning (ch. 5).
- Rumelhart, D. E., Hinton, G. E., & Williams, R. J. (1986). Learning representations by back-propagating errors. Nature, 323, 533–536.
- LeCun, Y., Bengio, Y., & Hinton, G. (2015). Deep learning. Nature, 521, 436–444.
- Nair, V., & Hinton, G. E. (2010). Rectified linear units improve restricted Boltzmann machines. Proceedings of ICML-10.
- Glorot, X., Bordes, A., & Bengio, Y. (2011). Deep sparse rectifier neural networks. Proceedings of AISTATS.