SDN Controllers Explained: ONOS and OpenDaylight Compared - 夜莺博客

SDN Controllers Explained: ONOS and OpenDaylight Compared

The first wave of software-defined networking promised that a central controller would replace the distributed control plane entirely. That did not happen - BGP, OSPF and IS-IS are still doing their jobs - but controllers did find durable niches: programming a fabric's overlay, managing a packet-optical domain, coordinating a segmented campus, and acting as the policy brain for networks where the business intent is the input rather than the device configuration. If you are evaluating one today, the two platforms worth understanding are ONOS and OpenDaylight.

What a Controller Actually Does

A controller is three things bolted together:

  • A southbound adapter layer. Drivers that speak OpenFlow, NETCONF, gNMI/OpenConfig, P4Runtime, SNMP or a vendor CLI, and normalise device state into the controller's model.
  • A state store and application core. The authoritative view of topology, links, hosts, intents and flows.
  • A northbound API. REST, RESTCONF or gRPC, used by applications and by your automation.

The value is in the middle layer. If you only need to push configuration, you do not need a controller - you need a configuration management system. A controller earns its place when something must be recomputed as the network changes.

Architecture: ONOS

ONOS was designed by and for service providers, and the architecture reflects that emphasis on scale and availability:

  • Distributed but logically centralised core. Every instance runs the full set of applications; state is partitioned across instances rather than sharded by function.
  • Consistent stores. Topology, device, link and flow state live in replicated stores with a strong consistency model, so any instance can answer any query.
  • Scale-out clustering. Instances join a cluster and agree on mastership of devices; mastership is per-device, which means a single device is not programmed by two pods at once.
  • Built for hot paths. Intent-based northbound APIs let an application express "connect A to B with this policy" and let the core compute the flows.
# a quick ONOS REST round trip
curl -s http://onos:8181/onos/v1/devices | jq '.devices[] | {id, available, type}'
curl -s http://onos:8181/onos/v1/links | jq '.links | length'
curl -s -u onos:rocks http://onos:8181/onos/v1/topology | jq .

Architecture: OpenDaylight

OpenDaylight takes a different route, closer to a model-driven service platform than a monolithic controller:

  • OSGi/Karaf runtime. Features are installed individually, so the running system contains exactly the plugins you asked for.
  • MD-SAL (model-driven service adaptation layer). YANG models define the data tree and the RPCs; every plugin and every northbound consumer talks to that tree rather than to bespoke APIs.
  • Strong NETCONF/RESTCONF story. Because the platform is YANG-first, it is a natural fit for environments where the devices are already modelled - which is most modern network operating systems.
  • Clustering via Akka. Scale-out is supported, but the consistency model and the operational complexity differ from ONOS, and in practice most OpenDaylight deployments are small clusters or single instances.

Side by Side

Dimension ONOS OpenDaylight
Origin and centre of gravity Service provider, transport and fabric Enterprise and vendor ecosystem
Data model Intent plus its own stores YANG / MD-SAL
Clustering maturity Designed for it; multi-instance is the normal mode Possible; smaller deployments are more common
Southbound strengths OpenFlow, NETCONF, P4Runtime, optical NETCONF/RESTCONF, OpenFlow, vendor plugins
Learning curve Application and intent concepts YANG, Karaf, feature management
Best fit Fabric and transport control, TE, optical Model-driven multi-vendor automation

Clustering Behaviour Is the Real Test

Both platforms claim scale-out. What matters operationally is what happens when a node dies:

  • Quorum loss. With an even number of instances and a partition, mastership decisions can stall. Size clusters odd and small - three instances handle most production workloads better than five badly-tuned ones.
  • Mastership stickiness. Prefer stable device-to-instance assignment. If mastership moves on every reconnect, flow programming churns and the network sees repeated transient reconvergence.
  • State recovery cost. After a full cluster restart the controller must relearn topology before it can program anything. Know how long that takes for your fleet size, and whether the network is stable during the gap.

Where Controllers Actually Get Used

  1. Data centre fabric underlay and overlay. A controller that owns the VXLAN tunnel endpoints and the tenant mapping can add a rack without a human touching a switch.
  2. Transport and optical. Path computation across a WDM or packet-optical layer is a genuine optimisation problem, and the controller is where the algorithm lives.
  3. Campus segmentation. Mapping a policy (group X may reach group Y) onto ACLs and VLANs across hundreds of switches is exactly the kind of recomputation a controller is good at.
  4. Service chaining. Steering traffic through a sequence of virtual functions, and re-steering when one moves, needs an entity that knows the whole picture.

Why Controllers Lost Mindshare

Three honest reasons:

  • Streaming telemetry replaced polling. A lot of what controllers were sold for - visibility - is now available directly from the devices via gNMI subscriptions, at higher frequency and lower cost. If that pipeline already exists, as described in this gNMI streaming telemetry guide, the controller's monitoring justification evaporates.
  • Declarative automation is simpler. For configuration that does not need recomputation, a source of truth plus a template engine plus an API is easier to operate than a controller cluster. The IPAM and source-of-truth pattern in this NetBox IPAM guide covers most of what teams actually needed from a controller's northbound API.
  • Routing protocols got better. Segment routing, EVPN and BGP-based fabrics do distributed path computation well enough that centralised control is only justified for TE or for policy.

Choosing

  • Transport or optical fabric with path optimisation? ONOS.
  • Multi-vendor, model-driven device automation where you already have YANG models? OpenDaylight.
  • Configuration push without recomputation? Neither - use an orchestrator, or a model-driven workflow against the device APIs, and see the declarative examples in this Terraform network automation guide and this SR OS model-driven configuration guide.
  • Small network? The controller is another cluster to run. Count the operational cost before the automation benefit.

Deployment Advice

Start with read-only. Point the controller at the devices, let it discover topology, and compare its view against your source of truth for a month. Discrepancies at this stage are free to fix; the same discrepancies discovered after the controller has write access are incidents. Then give it write access to one non-critical domain, and only then expand. The controller becomes a tier-zero dependency the moment it can change forwarding - monitor it, back up its database, and document the procedure for running the network if it is down.

原文链接:https://docs.onosproject.org/