P4 and BMv2: Programming the Network Data Plane - 夜莺博客

P4 and BMv2: Programming the Network Data Plane

Every switch you have configured defines behaviour in the same way: the vendors decided which protocols exist, and your job is to fill in the tables. P4 inverts that. You write the parser, the match-action pipeline and the deparser, and the target - hardware ASIC, software switch or DPU - compiles it. The result is that new protocols and new telemetry can be deployed at the speed of a software release rather than a hardware cycle.

This article covers the architecture, a complete working forwarding program, and how to run and test it on BMv2, the reference P4 software switch.

PISA: The Target Model

The Protocol Independent Switch Architecture describes a pipeline of stages, each containing a parser, a set of match-action tables and a deparser. In the v1model used by BMv2 the shape is:

                 +-------------------------  ingress  -------------------------+
packet -> parser | tables -> tables -> tables (match+action, single pass)     | -> deparser
                 +-----------------------------------------------------------+
                 +-------------------------  egress  --------------------------+
                 | tables -> tables -> tables                                  | -> deparser -> out
                 +-----------------------------------------------------------+

The constraints that shape every P4 program come from this model: tables cannot loop, each packet passes through the pipeline once, state must live in registers, counters or meters, and any operation not expressible in a table has to be handled by the control plane via digests or via a recirculation.

A Working L3 Forwarding Program

/* l3fwd.p4 - minimal IPv4 forwarding with a VLAN-aware parser */
#include <core.p4>
#include <v1model.p4>

typedef bit<48> macAddr_t;
typedef bit<32> ip4Addr_t;

header ethernet_t { macAddr_t dstAddr; macAddr_t srcAddr; bit<16> etherType; }
header ipv4_t     { bit<4> version; bit<4> ihl; bit<8> diffserv; bit<16> totalLen;
                    bit<16> identification; bit<3> flags; bit<13> fragOffset;
                    bit<8> ttl; bit<8> protocol; bit<16> hdrChecksum;
                    ip4Addr_t srcAddr; ip4Addr_t dstAddr; }

struct headers { ethernet_t ethernet; ipv4_t ipv4; }
struct metadata { }

parser MyParser(packet_in packet, out headers hdr, inout metadata meta,
                inout standard_metadata_t standard_metadata) {
    state start { transition parse_ethernet; }
    state parse_ethernet {
        packet.extract(hdr.ethernet);
        transition select(hdr.ethernet.etherType) {
            0x0800: parse_ipv4;
            default: accept;
        }
    }
    state parse_ipv4 { packet.extract(hdr.ipv4); transition accept; }
}

control MyIngress(inout headers hdr, inout metadata meta,
                  inout standard_metadata_t standard_metadata) {
    action drop() { mark_to_drop(standard_metadata); }
    action ipv4_forward(macAddr_t dstMac, bit<9> port) {
        hdr.ethernet.dstAddr = dstMac;
        hdr.ethernet.srcAddr = 0x001122334455;
        hdr.ipv4.ttl = hdr.ipv4.ttl - 1;
        standard_metadata.egress_spec = port;
    }
    table ipv4_lpm {
        key = { hdr.ipv4.dstAddr: lpm; }
        actions = { ipv4_forward; drop; NoAction; }
        size = 1024;
        default_action = drop();
    }
    apply {
        if (hdr.ipv4.isValid()) {
            if (hdr.ipv4.ttl <= 1) { drop(); }
            else { ipv4_lpm.apply(); }
        }
    }
}

control MyDeparser(packet_out packet, in headers hdr) {
    apply { packet.emit(hdr.ethernet); packet.emit(hdr.ipv4); }
}

V1Switch(MyParser(), MyVerifyChecksum(), MyIngress(), MyEgress(),
         MyComputeChecksum(), MyDeparser()) main;

Note what is not here: there is no ARP, no routing protocol, no MAC learning. A real data plane handles those either by punting to the control plane or by adding tables. That is the honest trade-off of data plane programmability - you get exactly the behaviour you specify, and you must specify all of it.

Compile and Run on BMv2

# install the toolchain (Docker images are the path of least resistance)
docker run -it --rm -v $(pwd):/work p4lang/p4c:latest \
  p4c-bm2-ss --p4v 16 --std p4-16 -o /work/l3fwd.json /work/l3fwd.p4

# start BMv2 with two virtual interfaces
sudo simple_switch --log-console -i 1@veth1 -i 2@veth2 l3fwd.json &

# or generate the whole topology with Mininet
sudo python3 run_exercise.py --topo topo.txt --switch l3fwd.json

The generated JSON is the compiled pipeline description. The runtime CLI connects to it and installs entries - this is the control plane, and in production it is a P4Runtime gRPC client rather than a manual CLI.

simple_switch_CLI --thrift-port 9090

RuntimeCmd: table_dump ipv4_lpm
RuntimeCmd: table_add ipv4_lpm ipv4_forward 10.0.1.1/32 => 00:00:00:00:01:01 1
RuntimeCmd: table_add ipv4_lpm ipv4_forward 10.0.2.1/32 => 00:00:00:00:02:02 2
RuntimeCmd: table_dump ipv4_lpm

# verify forwarding with a capture on the veth pair
sudo tcpdump -i veth1 -e -n icmp
ping -c 3 10.0.2.1

Adding State: Counters and Registers

Counters answer "how much of this traffic"; registers answer "what did I see recently" and are the basis of almost every custom in-network telemetry feature.

counter(bit<48>, DirectCounter) src_mac_pkts;
register(bit<48>)(1024) last_seen;

action record(macAddr_t src) {
    src_mac_pkts.count();
    last_seen.write(standard_metadata.ingress_port, src);
}
table mac_seen {
    key = { hdr.ethernet.srcAddr: exact; }
    actions = { record; }
    counters = src_mac_pkts;
    size = 1024;
}

Debugging

  • Learning log. simple_switch --log-console prints each packet's table hits, which is the fastest way to see whether a table simply had no entry.
  • Table dumps. table_dump before assuming the pipeline is wrong - most "P4 is broken" moments are empty tables.
  • Capture on both veths. A packet that disappears between veth1 and veth2 was dropped by mark_to_drop or by a missing egress specification.
  • Port constraints. egress_spec must be a valid port, and in BMv2 a wrong port number silently drops the packet.

Where P4 Actually Runs

Software targets are for development. Production targets today are programmable ASICs in data centre switches, SmartNICs and DPUs, and - with increasing frequency - the packet processing pipelines inside firewalls and load balancers. The toolchain is portable across them in principle; in practice each target has its own architecture file and its own externs, so a program written for BMv2 needs work to port. The DPU case is the most accessible entry point because you have a full Linux environment alongside the programmable pipeline; the architecture is described in this DPU and SmartNIC guide.

When P4 Is the Wrong Tool

  • Standard forwarding. A merchant silicon switch with EVPN and BGP does it better, cheaper and with support.
  • Anything needing loops or unbounded state. The pipeline model will not express it; you want a userspace stack. The vectorised model in this FD.io VPP guide is often the right answer instead.
  • Teams without a lab. You cannot debug a P4 program by reading the configuration. Budget for a containerised lab, and if you already run Containerlab for network testing, the workflows in this Containerlab with SR Linux guide transfer directly.

The realistic adoption path is: learn on BMv2 with Mininet, prove one useful feature (custom telemetry, an unusual encapsulation, a specific DDoS filter) on a software target, then evaluate whether a programmable hardware target is justified by that one feature's value.

原文链接:https://research.cec.sc.edu/cyberinfra/hands-tutorial-p4-programmable-data-planes-0