gNOI and gRIBi: Network Operations and Routing APIs - 夜莺博客

gNOI and gRIBi: Network Operations and Routing APIs

gNMI gets most of the attention because it is the protocol that streams state and pushes configuration. But a real automation workflow needs more than configuration: it needs to reboot a line card, verify a software image, fetch a file, run a ping, and - for controllers and service chains - inject forwarding entries directly. Those jobs belong to gNOI and gRIBi, and confusing them with gNMI is a common source of stalled projects.

This article explains what each interface is for, what services exist today, how to enable them on real platforms, and where they are the wrong tool.

Three Interfaces, Three Jobs

Interface Purpose Data model Typical use
gNMI Configuration and state via paths OpenConfig / vendor YANG Declarative config push, telemetry subscribe
gNOI Operational RPCs outside the config plane Proto services Reboot, image install, ping, file transfer, certs
gRIBi Program the forwarding table directly AFT (abstract forwarding table) Controller-driven FIB, service chaining, SRv6 steering

All three ride on gRPC over HTTP/2, so they share the same transport, TLS and authentication model. That is the practical benefit: one certificate, one connection pattern, one set of client libraries.

gNOI Services Worth Knowing

gNOI is a collection of proto-defined micro-services. The ones that show up in real automation:

  • System - Ping, Traceroute, Time, Reboot, RebootStatus, CancelReboot, SetPackage, SwitchControlProcessor, KillProcess.
  • OS - Activate, Verify, Install. This is the RPC family that makes maintenance automation safe, because Verify lets you confirm the image before committing to an activate.
  • Cert - LoadCertificate, Rotate, Revoke, GetCertificates. The right place to push short-lived certificates from an internal CA.
  • File - Get, Put, Remove, Stat. Useful for retrieving support bundles without opening an SSH session.
  • Containerz - deploy and manage containers on platforms that support third-party workloads.

Support is uneven. Cisco NX-OS, for example, documents System, OS and Cert support with specific limitations - RebootStatus backed by show version, CancelReboot mapping to reload cancel, and certificate rotation RPCs not implemented. Read the platform matrix before you design a workflow around an RPC; the proto existing is not the same as the target implementing it.

Enabling It

gNOI is not configured separately from gNMI on most platforms - it is a child of the same gRPC agent. On NX-OS:

feature grpc
! enable the gRPC agent (also exposes gNMI and gNOI)
grpc certificate <trustpoint>
grpc port 50051
grpc use-vrf management
show grpc status

On SR Linux or SR OS the same endpoint serves gNMI; check the platform documentation for whether gNOI is enabled by default and which RPCs are exported.

A Working gNOI Ping

The gnoi_client tool from the openconfig repository is the quickest way to exercise a device:

# build or fetch the reference client
git clone https://github.com/openconfig/gnoi.git
cd gnoi && make

# run a ping RPC against the target
./gnoi_client -target 10.0.0.1:50051 -insecure   -rpc System -system_rpc Ping   -json '{"destination":"10.0.0.2","count":5,"source":"10.0.0.1"}'

The response is a stream of PingResponse messages, each carrying the round-trip time and, if supported, the originating address. This is materially better than screen-scraping show output: the result is typed, and failures are gRPC status codes rather than free text you have to parse.

Pair gNOI with your structured telemetry work. If you already subscribe to state via gNMI streaming telemetry, the same client certificate and channel configuration can carry gNOI calls, which keeps the security model consistent.

gRIBi: Programming the FIB

gRIBi is different in kind. Instead of configuring routing protocols and letting the device compute a FIB, a controller computes the forwarding entries and installs them directly using the AFT model. The RPCs are small:

service gRIBI {
  rpc Modify(stream ModifyRequest) returns (stream ModifyResponse);
  rpc Get(GetRequest) returns (GetResponse);
  rpc Flush(FlushRequest) returns (FlushResponse);
}

A ModifyRequest carries one or more AFTOperations - add or delete of a NH, NHG (next-hop group), IPv4Entry, IPv6Entry or MPLSEntry - each tagged with a session-level election ID, a redundancy group and a persistence mode. That structure is what makes it safe to run concurrently with other controllers: entries from a session with a lower election ID are superseded rather than merged, and a session that dies can either leave its entries in place (DELETE persistence semantics) or be cleaned up automatically.

# conceptual AFT operation for a /32 over a two-member next-hop group
{
  "session": {"election_id": {"low": 1}, "id": 42},
  "operations": [
    {"add": {"nh":     {"index": 10, "next_hop": {"ip_address": {"v4": "10.1.1.1"}}}}},
    {"add": {"nhg":    {"index": 20, "next_hops": [{"nh_index": 10}]}}},
    {"add": {"ipv4":   {"prefix": "203.0.113.9/32",
                        "next_hop_group": {"index": 20},
                        "persistence": "FIB_AND_RIB"}}}
  ]
}

The election ID is the piece people get wrong. Two controllers both programming gRIBi with default election IDs will fight, and the symptoms are intermittent forwarding changes that correlate with whichever controller reconnected last. Always assign election IDs deliberately and treat them as part of the design.

Where gRIBi Is the Wrong Answer

  • Normal routing. If a protocol can compute the path, let it. gRIBi is for paths no protocol can express - service chain steering, anycast pinning, traffic-engineered overlays driven by an external SLO engine.
  • Heterogeneous fleets without a fallback. Not every platform implements every AFT entry type. Design a rollback path and keep the protocol-derived routes authoritative where possible.
  • Small sites. The operational cost of a controller with election-ID discipline is not worth it below a few hundred devices.

Practical Adoption Order

  1. Standardise on gNMI for config and telemetry, with a real YANG model workflow such as the one in this OpenConfig and pyang guide.
  2. Add gNOI System Ping and OS Verify to your change automation - they are low-risk and immediately reduce manual SSH work.
  3. Move image management to gNOI OS Install/Activate with Verify gating the commit.
  4. Only then evaluate gRIBi, and only for the specific steering problem it solves.

A reference implementation on a modern network OS makes the interface concrete; the SR Linux walkthrough in this SR Linux CLI and gNMI guide is a good lab target because the model is published and the RPC set is visible.

原文链接:https://github.com/openconfig/gnoi