Case study
.jpg)
Before automation, it was reasonable to assume that cost scales with demand, and headcount must grow to match. When ticket volume on a hyperscaler's global network team jumped 115% in just 6 months, Astreya used AI to help them clear the backlog and fix the data, with only a 6% increase in headcount.
On this client's optical and IP network, most faults trigger an alert and clear on their own before anyone investigates. By the time an engineer gets to it, the job has shifted from diagnosis to confirmation — prove the link is genuinely healthy, record what happened, close it out. That's what separates a transient event from the opening signature of a failure that hasn't finished yet. Skipping it defers a decision about whether something is actually wrong.
Astreya's operational support engineers do this across the client's global optical and IP infrastructure, with one goal: keep it available, at whatever volume comes through. For years the arithmetic held: the network grew, ticket load grew with it, staffing grew to match. The model was predictable because the ratio was.
That changed fast. Between January and June 2026, monthly ticket volume climbed from roughly 9,300 to more than 20,000 — per-engineer load roughly doubled in six months, on infrastructure that can't stop and wait for the team to catch up.
Two things broke.
As engineers ran out of hours, the checks that verify a cleared link went undone and tickets accumulated. A backlog on network infrastructure is corrosive because it buries the tickets that matter under a pile of tickets that don't.
Every resolved ticket requires a Root Cause Code, and the taxonomy was bloated: duplicate codes, unused codes, more options than anyone working at volume could navigate reliably. Measured against client-adjudicated ground truth, human root cause coding accuracy came in at 50%.
The conventional response to a curve like this is to hire against it. But this can be costly and unsustainable, especially for sudden spikes in demand.
So instead of focusing on increasing headcount to manage the queue, Astreya first focused on filtering the queue by asking which tickets needed human intervention at all.
Many didn’t. For instance, the tickets describing conditions that had already cleared didn't need judgment, but they did need availability — someone (or something) to look at, verify, and close them. That's a different requirement, and it's the one machines are good at.
Astreya built an AI agent that guides engineers through ticket resolution and closes tickets tied to ephemeral, self-resolving issues without human involvement.
The gating logic isn't time-based. It runs continuous checks on link status, and a ticket becomes eligible to close only when it passes every check. Anything that fails a check, or that the system can't resolve, routes to an engineer.
Astreya is direct about the limits here: passing every check isn't the same as certainty, and a proportion of issues recur after a clean check, a gap Astreya hasn't yet quantified. Which ticket types can be automated is determined empirically, not by category. The agent analyzes a class of tickets, attempts automation, and defines scope accordingly. Straightforward cases like configuration changes and status checks are prime candidates for automation. Anything requiring expert analysis stays with the engineers.
The agent executes changes, not just closures, but only from a pre-approved set. Authorization happens once, at the class level, so no individual change waits on a gate. A change that fails reverts to engineer review. And because a failed change doesn't partially apply, there's typically nothing to undo.
The agent reduced how many tickets a person had to handle, but it didn't make the root cause data usable. So Astreya built an AI recommendation engine to assign the codes instead. Against the same client-adjudicated standard that put human accuracy at 50%, Astreya’s model runs at 95%. As of July 2026, root cause codes are being applied to tickets at that quality for the first time, and support engineers no longer assign them by hand at all.
That jump from coin flip to near-certainty matters because this data feeds capacity planning, vendor accountability, and reliability reporting. A wrong code propagates into a forecast, a vendor conversation, an availability number.
At 20,000 tickets a month, auditing every code isn't possible — the same volume pressure that made manual coding untenable applies to manual review. Astreya's model is a 5% sample, roughly 1,000 tickets a month, with the first cycle running in August 2026, built to catch model drift if and when it happens.
What's left for engineers now is the harder work: complex fault isolation, genuine root cause investigation — the work that builds expertise instead of consuming it.
Combining Astreya’s AI agent and the coding engine frees up roughly ten engineers' worth of capacity a year, which Astreya reinvests in higher-value work.
Support cost normally tracks support demand — headcount scales with volume, and the only real variable is how efficiently you hire. If this were business as usual, the team would have needed to nearly double. Instead, ticket volume rose 115% between January and June while headcount rose just 6%. By July, automation was carrying the bulk of Tier 1 resolution and producing root cause data accurate enough to act on, not just report.
For anyone running a large support estate, the big takeaway is the interval: six months between volume spike and automation running in production.
Astreya kept things moving by tapping the same people who've been working the tickets to decide what gets automated. Bring in an automation team that's never touched the queue and you'll get the wrong things automated, or the right things automated with no data to trust once they're running.
To learn more about our Enterprise AI services, visit our capabilities page or contact an AI expert.