Amazon Web Services scientists published a post on October 1, 2026 describing a "graph-centric" agentic AI system that turns a telecom network into a continuously synchronized graph, or digital twin, and uses cascaded graph algorithms to identify the root cause of failures. The authors, AWS principal applied scientist Imen Grida Ben Yahya and senior solutions architect Nameet Dutia, said the approach was demonstrated with NTT DOCOMO at Mobile World Conference earlier this year.
The problem the authors describe
According to the post, finding the root cause of a network failure can take hours, and for complex multilayer failures remediation in traditional network operations centers averages four to five hours and can extend to days.
The authors said the bottleneck is cognitive overload rather than engineering expertise, because human operators cannot correlate hundreds of alarms, configuration files and telemetry data across thousands of nodes faster than customer impact accumulates.
They argued that traditional temporal correlation — inferring that alarm A caused alarm B because it came first — fails in complex topologies where failures propagate through parallel paths, polling intervals blur timing, and the true root cause may generate no alarm at all.
How the cascade works
The system represents the network as a graph whose vertices are devices with attributes and whose edges are connections, ingesting dependencies, live alarms and key performance indicators from multiple sources. Amazon calls this a digital twin of a physical or software-defined network.
Three stages then narrow the search space. Decomposition identifies the most-connected parts of the topology and localizes analysis to boundary nodes, cutting candidates from thousands of nodes to hundreds. Clustering with community detection algorithms — Louvain or label propagation — narrows candidates from hundreds to tens. Centrality ranking then scores the remaining candidates.
Crucially, the authors said the centrality measures are recomputed against the alarm set rather than the full graph. Their personalized PageRank variant seeds random walks from alarming nodes; their degree centrality counts only edges to alarming nodes; and alarm-relative closeness measures average distance to alarming nodes only. The agentic layer classifies each affected subgraph as hierarchical, star or mesh using metrics including degree distribution, diameter and hierarchical depth, then picks the algorithm combination.
The agentic layer
The post said agents first query an incident knowledge base and apply a prescribed remedy directly if an incoming alarm pattern matches a stored incident with high confidence; otherwise they invoke the full cascade after a complexity triage based on the number of affected nodes, their topological spread and the semantic clarity of the fault signature.
Once candidates are ranked, agents expand a failure subgraph to dependency neighbors two or three hops away and consult the alarm timeline, incident knowledge base and runbooks. The resulting determination carries a confidence score and either opens a trouble ticket or adds to an existing one, with operations-center feedback for continuous learning. An on-demand AI assistant lets engineers query the digital twin interactively.
What comes next, according to Amazon
The authors proposed two research directions. Under "graduated autonomy," an agent would move from advisory recommendations to supervised execution and eventually end-to-end remediation for well-understood, reversible, low-blast-radius incidents, with promotion tied to predefined evaluation criteria and shadow-mode results and autonomy reduced if performance drifts.
Under "self-learning agents," an agent would propose repeatable resolution patterns as candidate skills, but no skill would become active until explicitly approved by the user. The post also said graph neural networks and spatiotemporal deep learning could augment the deterministic cascade.
Key facts and where they come from
- Amazon said the approach was demonstrated with NTT DOCOMO at Mobile World Conference, achieving root cause analysis in minutes on commercial networks.
We demonstrated our approach with NTT DOCOMO at the Mobile World Conference (MWC) earlier this year, achieving root cause analysis in minutes on commercial networks.
- Traditional network operations centers average four to five hours to remediate complex multilayer failures, per the post.
For complex multilayer failures, remediation in traditional network operations centers (NOCs) averages four to five hours and can extend to days.
- Dependency graphs autogenerated from SDN and NFV controllers enabled Bayesian fault localization at 95% accuracy in under 30 seconds with no manual rules.
they enabled Bayesian fault localization at 95% accuracy in under 30 seconds with no manually authored rules
- The pipeline runs a three-stage cascade of graph algorithms, each narrowing the search space.
the system orchestrates a three-stage cascade of graph algorithms, each stage narrowing the search space for the next
- Clustering uses Louvain or label propagation community detection to cut candidates from hundreds to tens.
community detection algorithms (Louvain or label propagation) group nodes that frequently interact or share dependencies
- Agents first check an incident knowledge base before running the full cascade.
The first step is always to query the incident knowledge base: if the incoming alarm pattern matches a stored incident with high confidence, agents apply the prescribed remedy directly.
- Under the proposed self-learning design, no agent-proposed skill becomes active without explicit human approval.
The user retains full governance: no skill becomes active until explicitly approved.
