Telstra’s massive July outage was ultimately caused by a failure to recognise the importance of a critical network function, with an independent investigation finding a cascade of governance shortcomings behind the two-day disruption.

The investigation found Telstra had failed to treat Network Time Protocol (NTP) as a high-risk function of its network, with unclear ownership, inadequate technical expertise and weaknesses in documentation, change management and escalation processes.

Those failures allowed a known bug in an obsolete NTP server to go unfixed, despite warnings from the vendor and previous problems with the system.

The two-day outage – which started in the early hours of July 8 and affected 8.8 million phone users and interrupted Triple Zero calls, trains and payments into the next day – caused chaos nationwide.

Internal investigations confirmed that vendor warnings about the bug, which caused parts of the mobile network to reject connecting devices after the NTP server reset its clock back to 2009, had been repeatedly ignored by Telstra technicians who failed to treat it as a high-priority issue.

Telstra quickly contracted Technology Audit Partners (TAP) to initiate an independent investigation into the incident, with the newly released final report identifying a cascade of governance shortcomings that blew up when technical staff scrambled to fix the issue.

Review of the processes that led to the outage reached back to 2020, when Telstra introduced a new NTP timing chassis to replace its original hardware – but technicians failed to register that the new configuration had, TAP’s report found, “degraded the mobile core timing architecture.”

Last October, there were hints of the problems to come when an issue with the servers’ failover configuration caused problems – but that issue, TAP notes, “was observed but not noted and no problem ticket was created for further investigation”, pointing the way towards disaster in July.

During that incident, TAP found, Telstra’s IP Multimedia Subsystem (IMS) Core – which manages users’ access to services on the Telstra mobile network – “was significantly impacted… in such a way that access attempts… were blocked from using the network regardless of available capacity.”

A failure of governance

Although the outage was caused by a technical issue, TAP’s investigation concluded that the real problem was a failure of governance – meaning that Telstra should have had adequate documentation, investigation, escalation and repair processes in place to avoid and fix it.

“Unclear and insufficient” responsibility for NTP, it found, “and lack of recognition of its importance… led to weakness in this area and not treating it as a high-risk function of the network for the purpose of design, operations, resources and investment prioritisation.”

The impact of this lack of oversight was exacerbated by a lack of “architecture and design oversight” and the absence of an end-to-end architecture for mobile timing – stemming from Telstra’s failure to recognise NTP that is a Sovereign Function widely used across its network.

TAP also blamed poor change management, configuration control and incident management, noting that Telstra’s network team had “insufficient technical expertise for NTP” and flagging the “need for stronger competency and more curiosity to investigate when something looks wrong.”

Telstra, TAP found, should have increased “attention to the operation and management process discipline in all areas” – with the report offering 11 recommendations including stronger NTP ownership, architecture and design, configuration control, and change and problem management.

Despite its shortcomings, TAP noted that Telstra responded quickly to identify the issue – noting that “once the full Major Incident Management (MIM) process fully kicked in at 4:47am the process was organised and effective in one of the most complex recoveries we have ever seen.”

Yet because it hadn’t treated NTP as being important enough to require detailed documentation of changes made to the systems, TAP said, the service “appears to have fallen into the gap of coverage that impacted early outage detection and determination of cause”.

As they fought to identify and fix the problem, it said, this deficiency left the MIM teams struggling to find “information about the change… or the people to contact… or records of the servers themselves.”

Where to from here?

Telstra has implemented a range of changes since the incident, and accepted all findings of TAP’s investigation – which, Brady said, “provide valuable external insights into where we need to strengthen our network and management processes.”

The outage led Telstra’s board to penalise Brady and several other Telstra senior executives, with further penalties possible during fiscal 2027 since the outage occurred early in the current financial year.

The findings of TAP’s review “confirm what consumers have long suspected,” consumer telco advocate ACCAN said, noting that “Australia cannot keep relying on telcos’ ‘best efforts’ when it comes to the reliability of an essential service.”

“A function critical enough to bring down calls, payments and emergency services nationwide should never have been left without clear ownership,” ACCAN CEO Carol Bennett noted in offering the incident as evidence that the industry needs “clear and enforceable reliability standards.”

“It is deeply disturbing,” she added, “that a business as profitable as Telstra did not have the in-house capability to properly maintain a function this critical to its network.”

As management consultants dissect the incident – Melbourne Business School professor of leadership Will Harvey called it “important lesson in leading under pressure” and noted that “apologies are remembered for days, but actions are judged for years” – Brady is looking forward.

“This outage should not have happened,” she wrote while committing to “use what we have learned… from this outage to not only address what went wrong, but to make Telstra stronger and our services even more resilient and reliable for our customers.”

“Modern networks are complex, but complexity is not an excuse.”