Topology Discovery and Reconciliation in Carrier OSS
Real-time topology accuracy prevents orders from failing when network records drift out of sync.

A fiber provider runs a serviceability check on a prospective customer's address, and the OSS says yes, capacity is available. Multiply that by hundreds of orders a month: the pattern is topology drift, the condition where what the OSS believes about the network and what the network actually contains have become two different records, not one record with a lag.
Why topology data in carrier OSS diverges from network reality
The failure above traces back to a basic fact about how carrier OSS platforms hold data. There's no single file anyone can point to and call the authoritative one. When three systems each hold a partial, slightly different account of the same physical thing, the organization has no ground truth to check against: only competing claims.
It's tempting to treat this as a tooling gap, something a better discovery platform could close by polling the network more often or more precisely. That framing misses where the problem actually sits. If the model has no way to hold one canonical identity for a given network element, running discovery jobs more frequently just produces more frequent copies of an already-fractured picture. The constraint is structural, a matter of the data model rather than polling interval.
GDi's Ensemble OSS Discovery and Reconciliation product describes its core job as unifying and persisting network data in a normalized format, then quickly identifying discrepancies between what's downloaded from the network and what's stored in inventory. The word "quickly" matters here. Speed in spotting discrepancies only has value once the data has been normalized into one format first. Without that normalization step, a faster comparison just returns a faster mismatch; it tells an operator nothing has reconciled, only that the two records still disagree.
This cost appears in ordinary operational decisions, not dramatic outages. If a port filled or a segment of plant went out of service overnight, the qualification is wrong before the customer has even signed anything. That produces incorrect availability answers, orders that get accepted and then fail downstream, and truck rolls to addresses that were never actually serviceable. None of this requires a network outage or an equipment failure. It only requires the OSS record and the network to have quietly stopped agreeing with each other.
How legacy OSS architecture makes the divergence permanent
Topology drift in most carrier environments is the predictable output of how legacy OSS platforms were built. Two structural features do most of the work: a monolithic design built for an earlier generation of services, and a hard dividing wall between OSS and BSS that was never meant to be crossed in real time.
Start with the wall. OSS, in the traditional carrier stack, handles the network: inventory, provisioning, fault management, performance monitoring. Every handoff through custom code is a point where some context about a network element, a status, a relationship to another element, can get dropped or mistranslated, and those small losses accumulate into the kind of mismatch a provisioning team discovers only when an order fails.
The monolithic design compounds the problem. That design assumption made sense in its era. Reconfiguring a monolithic, on-premises OSS to track that kind of change usually requires custom development work for each new service pattern. The system is perpetually behind the network it's supposed to describe.
The clearest expression of both problems together is the batch-reconciliation model. Legacy systems typically defer reconciliation to scheduled jobs, often run overnight, so the OSS's understanding of the network is, by design, always some number of hours old. That was an acceptable trade-off when provisioning cycles ran in days and service changes were rare. It stops being acceptable once a provider wants any form of closed-loop automation, because a closed loop that acts on overnight-old data is making decisions about a network state that may no longer exist by the time it acts.
What real-time reconciliation requires at the data model level
Solving for real-time reconciliation isn't a matter of running the same batch job more frequently. A faster batch job is still a batch job: it still depends on moving data between systems that disagree about its shape. Real reconciliation in real time requires a different data model altogether, one in which a single normalized identity for every network element is held and used simultaneously by every operational function that touches it, rather than copied and reconciled after the fact.
This is where it's worth looking closely at what ETL, Extract, Transform, Load, actually implies about a system's architecture. Reconciliation built on ETL is managing the consequences of drift, not removing the condition that causes it. GDi's Ensemble platform, for instance, uses an ETL process to export Resource Inventory data into unified structures for comparison, a capable and well-engineered implementation of this approach. But comparison after transformation is still reconciliation after the fact. It catches divergence once it has already happened instead of preventing the divergence from forming.
A genuinely unified data model works differently. And it needs to let qualification, provisioning, and activation functions read and write the same record directly, without a translation step between them. Under that design, a change made during provisioning is visible to activation immediately, because there was never a second copy of the record to begin with. Changes propagate at the moment they happen rather than waiting for the next reconciliation cycle to surface them.
An industry framework for telecom software architecture points toward this outcome at the industry level. ODA groups OSS and BSS capabilities into modular software components that deploy on a shared ODA Canvas, replacing the traditional OSS/BSS split with a neutral framework that doesn't care which side of the old wall a given function used to sit on. The significance of that design choice is architectural: it dissolves the dividing wall between network-facing and commercial-facing systems rather than building a better bridge across it.
Where the Gap Becomes Visible: FTTH, Dedicated Internet, and Carrier Ethernet Delivery
Topology drift stays abstract right up until it reaches an actual order. In fiber-to-the-home delivery, two workflows depend entirely on topology accuracy: the feasibility check run at the point of sale, and the conditional appointment reservation made for installation. A feasibility check depends on address inventory being correct down to the individual premise. An appointment reservation depends on knowing, at the moment the appointment is booked, whether plant capacity is actually available to support it. Both of these break the instant the topology record the OSS holds diverges from what's physically been built or consumed in the field.
Dedicated internet access shows the gap differently, through the chain running from qualification to activation. A misrepresented port assignment, a logical link that was never recorded, or a capacity figure that's gone stale all produce the same outcome: a provisioning failure that someone has to manually investigate before the order can move forward again. That manual diagnosis step is pure overhead, time spent because the record of the network and the network itself had already parted ways before the order was placed.
Carrier Ethernet delivery follows the same pattern, and the common thread across all three service types is where errors actually accumulate: at the handoffs. At the seams, the architectural fragmentation described earlier becomes a specific failed order with a specific customer waiting on it.
Why AI agents cannot compensate for a fragmented topology record
Adding AI agents to an OSS that already has fragmented topology data doesn't correct the fragmentation. It makes the consequences worse, because an agent making decisions and taking action needs coherent read and write access across the full operational picture to act reliably, and a fragmented record cannot supply that.
Much of the AI deployed in telecom over the past several years has worked as an overlay: machine learning models bolted onto legacy systems to handle predictive maintenance or anomaly detection without touching the underlying data architecture. An agent operating across multiple schemas has to translate between them at every handoff, and each translation is another place where errors accumulate and context gets lost, the same failure mode described earlier in the OSS/BSS integration layer, just automated and running faster.
The sharpest version of this risk is what happens when point-solution AI agents get layered onto an OSS that's only partially consolidated. Each agent optimizes for whatever objective it was built to serve, working against whatever slice of data it has access to, and the result can be measurably worse than doing nothing automated. That outcome is the predictable result of giving autonomous decision-makers partial, disagreeing views of reality and asking them to act on it anyway.
Making the data model coherent before asking agents to operate on it, and keeping agents and human operators working inside the same governed system rather than separate ones, resolves this. That means the same APIs, the same audit logs, the same permission structures, so every action an agent takes is visible and reviewable the same way a human operator's action would be. The orchestration layer that makes it safe to move from insight to action is an intent-based layer underneath the agents, a deterministic, model-driven framework that keeps automation predictable and auditable even as the agents themselves get more capable and more autonomous over time.
Why patching legacy OSS to be AI-ready does not resolve the structural problem
The strongest argument against consolidating an OSS onto a unified data model is economic reality. OSS modernization is itself a multi-year capital commitment, arriving at the same moment fiber capex is already consuming most of the balance sheet. The conditions that make consolidation necessary, aging architecture, rising service complexity, growing AI ambitions, are the same conditions that make consolidation hardest to fund and schedule.
TM Forum's Catalyst program offers the most credible evidence that incremental modernization is viable, demonstrating event-driven architecture adoption that doesn't require operators to walk away from decades of existing OSS/BSS investment. TM Forum's own Catalyst FAQ describes these as rapid-fire proof-of-concept projects, which makes the program a directional signal worth watching rather than a validated migration path operators can adopt wholesale today.
Even granting everything the Catalyst program demonstrates, event-driven architecture and modular ODA components lower the cost of migration relative to a full replacement, but they don't remove the underlying requirement for a unified data model. When that step is skipped, the integration debt described earlier doesn't disappear: it just relocates into whatever new middleware connects the modernized pieces to the parts still running the old way.
The operators best positioned to avoid this are the ones building without legacy architecture to carry forward. Greenfield operators building AI-native foundations from the start will move faster than incumbents managing partial consolidation, and that gap compounds with every new service type layered on top of a topology record that's still fragmented. CLECs and fiber operators face this pressure most acutely right now: the most commercially obvious FTTH footprints have largely been built out already, construction costs keep climbing, and competition from cable and fixed wireless is intense enough that margins have little room left to absorb the internal waste a fragmented OSS stack generates in failed orders, manual reconciliation work, and truck rolls that shouldn't have happened. For these operators, the real question is when to consolidate the topology record onto a single data model, and every quarter of delay adds more AI investment built on top of fragmented data that will eventually need to be rebuilt once the foundation underneath it finally changes.
A Topology-Aware, AI-Native OSS Architecture in Practice
A topology-aware, AI-native OSS consolidates the entire service delivery lifecycle, qualification, design, provisioning, and activation, onto one data model, so every function involved reads and writes the same canonical topology record directly, with no translation step standing between them.
In practice, that means service qualification, provisioning, dispatch, field execution, network activation, and billing connect to each other without manual reconciliation or scheduled system-to-system exports standing in the way. Discovery and reconciliation stop being a separate inventory process that gets synchronized against the real systems after the fact; they operate against the exact same data model that provisioning and activation already use. Configuration drift stops being a risk to manage, because there's no secondary system left for the primary one to drift away from.
Governed AI fits into this picture as a participant inside the same framework human operators already use. Compliance follows directly from this design: workflows and configurations get checked continuously against regulatory requirements, deviations get flagged before they become violations, and audit-ready documentation gets generated automatically, all because governed AI and human operators are working against the same data and the same permission model rather than separate ones that need to be reconciled against each other later.
TM Forum's ODA Conformance Test Kit gives vendors a concrete, programmatic way to verify that a given software component is genuinely plug-and-play ready for this kind of architecture, which lowers the cost of building toward it for application developers rather than leaving each operator to verify compatibility from scratch. Treated this way, OSS stops functioning as a back-office cost center to be minimized and starts functioning as the strategic platform that determines how fast an operator can actually sell, build, and activate the services its network was built to deliver.
Sources
- Generic discovery for computer networks
- Data Engineering Patterns for Cross-System Reconciliation in Regulated Enterprises: Architecture, Anomaly Detection, and Governance
- Update and procurement of telecom lines using automated reconciliation of device information
- Update and procurement of telecom lines using automated reconciliation of device information
- Background discovery agent orchestration
- The transformative impact of AI and generative AI on OSS and BSS in telecommunications
- Ethernet Topology Discovery: A Survey
- Carrier ethernet service discovery, correlation methodology and apparatus


