---
title: "Amazon Pushes A.I. Agents That Use Network Graphs to Find Failures in Minutes"
description: "Company researchers describe a three-stage pipeline, demonstrated with NTT Docomo, that aims to cut fault-finding from hours to minutes"
author: "rews desk"
published: 2026-10-01T12:54:31.644Z
modified: 2026-10-02T01:16:57Z
url: https://rews.cc/a/amazon-pushes-a-i-agents-that-use-network-graphs-to-find-fai-5a9548
language: en
tags: ["ai", "networks", "automation", "gpt", "ntt", "tech"]
publisher: "Rews (https://rews.cc)"
---

# Amazon Pushes A.I. Agents That Use Network Graphs to Find Failures in Minutes

*Company researchers describe a three-stage pipeline, demonstrated with NTT Docomo, that aims to cut fault-finding from hours to minutes*

By rews desk · October 1, 2026 · https://rews.cc/a/amazon-pushes-a-i-agents-that-use-network-graphs-to-find-fai-5a9548

## In brief

- Amazon researchers published a post advocating graph-based A.I. agents for locating network failures
- Remediation of complex failures in traditional operations centers averages four to five hours and can take days
- A three-stage cascade narrows candidate root causes from thousands of nodes to a ranked handful
- Treating networks solely as correlation problems fails when the true cause sets off no alarm at all
- Amazon says it demonstrated the approach with NTT Docomo at MWC, finding root causes in minutes

Amazon researchers say they have built an approach that lets A.I. agents find the root cause of network failures in minutes rather than hours, detailing the system in a company research post published on Oct. 1.

The problem they are attacking is slow fault-finding in big networks. For complex multilayer failures, remediation in traditional network operations centers averages four to five hours and can stretch to days, according to the post. The bottleneck, the authors argue, is not engineering expertise but human overload: no operator can correlate hundreds of alarms, configuration files and telemetry across thousands of nodes faster than customers feel the damage.

Older tools make it worse. Many systems lean on temporal correlation: if alarm A precedes alarm B, the machine infers that A caused B. That heuristic fails in complex topologies where failures travel down several parallel paths, the intervals between system status checks muddy the timing, and the true root cause may set off no alarm at all.

The Amazon team’s answer starts by treating the network as what it natively is: a graph. Every device, link and service dependency becomes a vertex or an edge in a continuously synchronized “digital twin” that ingests network dependencies, live alarms and performance data from across segments and layers. Against that graph, the system runs a three-stage cascade, each stage narrowing the search for the next.

First, decomposition identifies the most-connected parts of the topology, cutting the candidate set from thousands of nodes to hundreds. Then community-detection algorithms, Louvain or label propagation, group nodes that frequently interact or share dependencies, paring the field to tens. Last, a suite of centrality algorithms ranks the survivors by how likely each is to be the root cause.

The twist is that every centrality measure is recomputed against the failure, not the full graph. The group’s version of personalized PageRank seeds its random walk from the nodes issuing alarms and traces faults upward, which suits hierarchical layouts. Its degree-centrality algorithm counts only edges to alarming nodes, so a gateway with 50 connections but none to an alarming node scores zero, while a switch with five alarm connections scores five. A third measure ranks the node closest to the failure cluster. An agentic layer classifies each broken subgraph as hierarchical, star or mesh and picks the combination of algorithms accordingly.

The post traces a lineage of graph methods in networking, from early topology maps to knowledge graphs that added semantics, to alarm-correlation graphs that condensed tens of thousands of raw alarms into a single causal chain within minutes. Later, dependency graphs generated automatically from software-defined networking and network functions virtualization controllers enabled Bayesian fault localization at 95 percent accuracy in under 30 seconds, with no manually authored rules. Graph neural networks went further, learning dependencies that topology alone could not show.

The researchers said they demonstrated the system with the Japanese carrier NTT Docomo at the MWC telecom trade show earlier this year, achieving root cause analysis in minutes on commercial networks. The post, by Imen Grida Ben Yahya and Nameet Dutia, is Amazon’s own account of its work; the results have not been independently reviewed.
