A07 · ITAI 1370 · Deep Learning Through Simulation

Predicting HVAC Failures
with a Neural Network

A feed-forward network that learns whether a commercial rooftop unit will fail within 30 days — watch it think, watch it get one wrong, and watch backpropagation teach it better.

▼ SCROLL TO EXPLORE
The Problem

Rooftop units fail without warning. The data that predicts it already exists.

An unexpected RTU failure at a commercial site means emergency labor rates, downtime, and lost product. Every routine service call already records the features that predict failure — age, maintenance recency, heat load, fault history. A neural network can learn the pattern.

1Supervised Learning

The network trains on labeled history: past service calls where the outcome — failed or didn't — is already known. It learns a function from features to outcomes (Goodfellow et al. 2016, ch. 5).

2The Architecture

4 input features → Hidden Layer 1 (4 neurons, ReLU) → Hidden Layer 2 (3 neurons, ReLU) → 2 output classes. Small enough to see every connection; deep enough to learn feature interactions.

3Why Depth Matters

Layer 1 detects simple feature combinations; layer 2 composes them into concepts like "wear & neglect" or "thermal stress" — hierarchical representation learning (LeCun, Bengio & Hinton 2015).

The Dataset

Five labeled service-call records

Illustrative data patterned on real commercial RTU field history. Each row is a training example: a feature vector paired with a ground-truth label.

UnitAge (yrs)Days Since ServiceAvg Ambient (°F)Fault Codes (90d)Outcome
RTU-0112210984FAILED
RTU-02345910DID NOT FAIL
RTU-0381201032FAILED
RTU-04 ★1530951DID NOT FAIL
RTU-056300993FAILED

★ RTU-04 is the trap: a very old unit that was just serviced. Age screams risk, recency says safe. Click the row — let's push it through the network.

Live Simulation · The Signature Moment

Run the network yourself

Line thickness = connection weight |w|. Fire the forward pass to see the prediction, reveal what actually happened, then apply backpropagation and watch the weights re-learn in real time.

RTU-04 loaded: age 15 yrs · 30 days since service · 95 °F ambient · 1 fault code. The network's initial weights over-trust unit age (see the thick orange-tinted lines). Run the forward pass.
THEORY · FORWARD PROPAGATION & WEIGHTED SUMS

Each neuron computes z = Σ(wᵢ·xᵢ) + b, then applies ReLU: f(z) = max(0, z). ReLU was chosen over sigmoid because it avoids saturating gradients, keeping the error signal strong for the learning step — a deliberate, theory-rooted design decision.

Goodfellow, Bengio & Courville (2016), ch. 6 · Nair & Hinton (2010), ICML · Glorot, Bordes & Bengio (2011)
Feedback & Learning

The error is the fuel

The network predicted WILL FAIL (74%). RTU-04 did not fail. That 0.74 error propagates backward through every connection — and each weight learns exactly its share of the blame.

▼ Age → h1.1 · WEAKENED

Gradient attribution assigns this connection the largest blame: it drove a confident wrong answer. Its weight shrinks proportionally to ∂Error/∂w.

▲ Service Recency → h1.2 · STRENGTHENED

Recent service was the true protective signal the network under-weighted. Its gradient pushes the weight up.

▼ "Wear & Neglect" → FAIL · MODERATED

The concept neuron stays useful, but its vote toward failure is tempered so age alone can no longer dominate the verdict.

THEORY · BACKPROPAGATION + GRADIENT DESCENT

Backpropagation uses the chain rule to compute ∂Error/∂w for every weight; gradient descent then updates each one: w ← w − η·∂E/∂w. The learning rate η controls step size — too large overshoots, too small crawls. One example nudges the weights; over many examples the pattern converges. After learning, re-running RTU-04 yields P(fail) ≈ 0.41 → the prediction flips to WILL NOT FAIL. The network learned that a freshly-serviced old unit ≠ a neglected old unit — an interaction no single weight could express.

Rumelhart, Hinton & Williams (1986), Nature 323, 533–536 · Goodfellow et al. (2016), ch. 6.5 & 8
Discussion Points

What this simulation demonstrates

Q1Theory → connection strengths

Initial weights encoded domain priors; then weighted-sum theory defined influence, ReLU kept gradients alive, and every strengthen/weaken decision after the miss followed gradient attribution — not intuition.

Q2What the error taught us

A 74%-confident miss exposed over-reliance on one feature. Backprop converted a single error into targeted, proportional updates. Next: tune η, add training data, hold out a validation set against overfitting.

Q3Real-world applications

Predictive maintenance on RTUs & chillers · cleanroom environmental-monitoring anomaly detection (pharma/biotech) · smart-building BAS optimization · live dispatch prioritization from fault-code telemetry.

References

Theoretical foundations

  1. Goodfellow, I., Bengio, Y., & Courville, A. (2016). Deep Learning. MIT Press. deeplearningbook.org — feedforward networks (ch. 6), training (ch. 8), supervised learning (ch. 5).
  2. Rumelhart, D. E., Hinton, G. E., & Williams, R. J. (1986). Learning representations by back-propagating errors. Nature, 323, 533–536.
  3. LeCun, Y., Bengio, Y., & Hinton, G. (2015). Deep learning. Nature, 521, 436–444.
  4. Nair, V., & Hinton, G. E. (2010). Rectified linear units improve restricted Boltzmann machines. Proceedings of ICML-10.
  5. Glorot, X., Bordes, A., & Bengio, Y. (2011). Deep sparse rectifier neural networks. Proceedings of AISTATS.