CASE STUDY

HVAC and Field Services.

A First Strategy case study.

Company name is held in confidence.

PDF Download the full case study (PDF).

The story

HVAC and Field Services

The diagnosis everyone agreed on was wrong

A multi-location HVAC and field services operator brought us in to answer one question: where AI actually fits in a business like theirs. Three vendors before us had pitched scheduling, the demos had impressed, and every pilot had died once dispatchers tried to use it. The Director of Operations stopped asking for a scheduling tool and started asking a better question.

What was at stake

Margin was bleeding through a back office nobody had flagged as the problem. Invoices were going out wrong, billing was running days behind the work, and the field-to-office handoff was eating hours a day in rework. Another miss on AI would lose the field's trust for any next attempt, and the operator had no appetite for a multi-year rebuild that might not work. The next move had to be small enough to be reversible and clear enough to prove out in weeks, not quarters.

What the floor revealed

We asked for a day on the floor before proposing anything. The morning with leadership lined up with the inherited diagnosis: dispatch was the bottleneck, scheduling was the move. The rest of the day did not.

The lead dispatcher was not scheduling in the dispatch software. It sat minimized on her screen while she worked from a hand color-coded spreadsheet refined over years. The software knew where technicians were; it did not know which one runs long on certain days and why, or which customers always add work once the truck is in the driveway. "This is how I schedule," she said. "The software is for billing." Automating her would have ripped out the part of the operation that worked.

The real friction sat downstream. At the billing desk we watched a clerk pull a paper ticket from a basket, squint at handwriting, guess at a part number, and dial a technician who would not pick up. By late morning, three invoices were stuck waiting for a callback and two had already gone out with errors. This was a normal morning. More than a third of invoices came back needing correction. Nearly half of jobs triggered a clarification call. Billing ran three days behind the work. About six hours a day, per location, went to cleaning up a handoff that should have taken seconds. When we traced the errors back, roughly eighty percent originated at one step: the paper form in the technician's hand. Not scheduling. A piece of paper.

The strategic shift

Instead of automating where every vendor wanted to sell, we fixed the handoff that was actually carrying the loss. The dispatcher's judgment stayed. The paper-first handoff at the point of service went. The trade-off was visible: a smaller first move than leadership had been pitched, but one we could prove in weeks instead of years, and reverse without harm if it did not work.

The second shift was how we would build. A confident team would have started construction. We ran cheap experiments first, designed to kill our own ideas before they got expensive. Evidence on the floor would force the next decision, not the inherited theory and not our own.

How we built it

Three experiments came first, days each. A scheduling visibility board in front of technicians lasted a week before becoming a place to post a fantasy football league; nobody used it, because nobody returns to a board when their next job arrives by text. A digital form was slower than paper and failed the moment a technician lost signal in a basement. The third experiment kept the paper, photographed each finished ticket, let AI read it into billing, and texted the technician to confirm anything unclear. Invoice errors fell from thirty-eight percent to nine with technicians doing nothing different. The signal was clear: capture had to be lighter than paper, not heavier.

The hard line for the build was fifteen seconds of capture in a customer's driveway, eight the target. Voice was the primary path, photographed paper the fallback, the same AI reading both. The first field test taught the lesson no specification would have: the system parsed a complicated part number cleanly, then returned pure garbage on a commercial rooftop, defeated by wind across the microphone. The team retrained on real field audio, compressors and truck engines and wind, instead of recordings from a quiet office.

During the pilot the billing manager reviewed every submission. One afternoon she stopped on a routine water-heater swap: a standard replacement, about two hours, that the AI had logged as eight. She pulled the photo, confirmed the job, screenshotted the error, and walked it to the developer's desk. The system was working as designed: not an AI that is never wrong, but an AI whose mistakes a human sees before a customer does. Her error log became the training data for the next week and the checklist the next location's reviewer would use.

We expanded one location at a time. The second, residential like the first, went smoothly. The third broke on purpose. It was fully commercial, where jobs run across days instead of hours, and the system had learned a false rule from residential: that a capture means a finished job. Errors climbed to fifty-five percent. We pulled it within days, put that location back on paper, and spent three days watching commercial crews. Commercial jobs have milestones; billing should fire on a milestone, not on every capture. The fix surfaced a constraint nobody had flagged, a billing system that did not support milestone billing, which forced an intermediary layer and a slipped timeline. The team owned the missed estimate directly with the client. The commercial version brought that location from fifty-five percent down to eleven, and the rollout finished, two locations to a wave.

Stable was not the same as governed. Once the system was live everywhere, extraction accuracy on plumbing jobs drifted while the headline dashboard stayed green. Catching it took watching disputes by service type, not by total. The team retrained on balanced data, added an alert whenever any segment fell more than five percent from baseline, credited the customers billed wrong during the drift window, and built three tiers of human oversight by job risk. A job type earned less oversight only on evidence: ninety days above ninety-five percent accuracy and zero boundary violations. Routine residential earned it. Nothing earned it by decree.

What changed

Across the operation, recovered billing reached about two hundred forty-seven thousand dollars a month, on the order of three million dollars a year. One location alone recovered about forty-seven thousand dollars in a single month, billing that would otherwise have been disputed, delayed, or written off.

Billing went from a three-day lag to same-day. The billing team's job inverted. Clerks who had spent their mornings deciphering handwriting became exception handlers, catching what the AI missed and feeding each correction back so it missed less the next time. Nobody was laid off. The same team now carries more locations than it could before. Capacity grew where headcount would have.

Invoice errors settled at nine percent on residential work and twelve on commercial, down from thirty-eight. Capture averaged about eleven seconds, faster than paper. Paper usage fell by roughly eighty percent, with the remainder kept as the fallback path. Clarification calls landed near zero.

What stands as proof

The Day One Audit recorded the baseline before any build. The pre-build arithmetic and the measured result agreed, which is the point: the number was earned, not promised. The pilot held in one residential location, then a second, then broke on purpose in a commercial context and was rebuilt to fit. The Audit and the Playbook stand as the engagement record.

We then handed the capability over the way we built it. We led the first project and built it ourselves. We handed the second to their team, trained a guide on their side, co-guided alongside them, and held the technical calls. We oversaw the third while they led it outright. They now run the process on their own, with a monthly advisory check-in.

The point was never to be permanently needed. It was to fix the right problem first, prove the approach in the operator's own building, and build the muscle to find the next one. The team knows the first problem will not be the last. They now know how to find it.

Where each of these started.

Every one of these engagements started with a day. A fixed-fee day in the business with leadership. Real work, not slides. A playbook within two weeks. Then a decision.

Start with a Day One

How it starts.

A day.

A fixed-fee day in your business with your leadership. Real work, not slides. Two weeks later you have a playbook, yours to run with us or without us. Day One.

A build.

You know what you want built. Tell us what it is. Inquire.

A team.

The systems are there and your people are not using them. We start with the work they actually do. Inquire.