PTP vs NTP Explained: Precision Time for Networks and Servers - 夜莺博客

PTP vs NTP Explained: Precision Time for Networks and Servers

NTP keeps a data centre consistent in milliseconds; PTP keeps a radio network consistent in nanoseconds, and mixing up the two requirements is an expensive mistake in either direction. This article explains what actually differs between the protocols - error sources, timestamping, and the clock roles that exist only in PTP - then shows a practical design for each, including the boundary versus transparent clock question that decides how a switch participates.

The fundamental difference

Property NTP PTP (IEEE 1588)
Typical accuracy 1-10 ms over WAN, <1 ms on LAN sub-microsecond to nanoseconds
Timestamping Mostly software Hardware (PHY/MAC) required for full accuracy
Transport UDP 123, unicast or broadcast L2 Ethernet or UDP 319/320, multicast by default
Hierarchy Stratum 0-15 Grandmaster, boundary, ordinary, transparent clocks
Best fit Servers, logs, DNS, K8s, general IT Telecom, financial trading, power, AV, industrial control, 5G

The accuracy gap comes from where the timestamp is taken. NTP timestamps in software, so the OS scheduler, interrupt handling and NIC queueing all inject jitter. PTP timestamps in hardware at the PHY, and the switches in the path correct for the time each packet spends inside them.

Clock types you will configure

  • Grandmaster (GM) - the source of time in the PTP domain, usually fed by GNSS or a PRTC.
  • Ordinary clock - a PTP endpoint with a single port, either master or slave.
  • Boundary clock - multi-port device that is a slave on one port and a master on the others, effectively regenerating time per segment. This is what a switch does when it terminates PTP rather than passing it through.
  • Transparent clock - does not become master or slave; it forwards PTP messages and adds a correction field for the residence time inside the device. Simpler, cheaper, and enough for most two-step topologies.

Junos-style configuration sketch

set protocols ptp clock-mode boundary
set protocols ptp domain 24
set protocols ptp slave interface ge-0/0/0.0
set protocols ptp master interface ge-0/0/1.0
set protocols ptp master interface ge-0/0/1.0 priority 128
set protocols ptp master interface ge-0/0/1.0 unicast-mode on
set protocols ptp master interface ge-0/0/1.0 unicast-mode transport ipv4
set protocols ptp master interface ge-0/0/1.0 unicast-mode clock-source 10.0.0.1 local-ip-address 10.0.0.2
set protocols ptp slave interface ge-0/0/0.0 unicast-mode on
set protocols ptp slave interface ge-0/0/0.0 unicast-mode transport ipv4
set protocols ptp slave interface ge-0/0/0.0 unicast-mode clock-source 10.0.0.1 local-ip-address 10.0.0.2
set protocols ptp slave interface ge-0/0/0.0 phy-timestamping
run show ptp clock
run show ptp parent
run show ptp port

The domain number must be identical on every device in the domain - a mismatch is the most common "PTP is not locking" cause. Keep one domain for production and a second for lab/testing so a lab grandmaster can never hijack production slaves.

Server side: PTP and NTP together

# linuxptp with hardware timestamping
ethtool -T ens1f0
ptp4l -i ens1f0 -m -s --summary_interval 0
phc2sys -s ens1f0 -c CLOCK_REALTIME -w -m -O 0

# keep NTP for everything that does not need nanoseconds
chronyc tracking
chronyc sources -v
timedatectl status

A common and correct pattern is stratified: grandmaster feeds PTP to devices that need it, the same reference feeds chrony or ntpd for the general estate, and PTP-derived system time is never mixed into an NTP hierarchy loop. Verify with phc_ctl, pmc and chronyc tracking before trusting anything.

Design checklist

Decide the required accuracy first, then work backwards: if milliseconds are enough, do not build PTP - NTP with a solid local stratum-1 or GPS source and clean anycast configuration is cheaper and far easier to operate. If nanoseconds are required, plan for hardware timestamping on every hop, a documented domain and priority scheme, redundant grandmasters with the same domain, and monitoring on the PTP port state rather than only on the receiving application. Related: chrony time sync and drift troubleshooting, Cisco IOS NTP configuration and fabric-wide PTP with Arista AVD.

Where the error actually comes from

The protocols are not just "the same thing at different speeds". They fail for different reasons, and knowing the error budget is what tells you whether a design will hold.

Error source NTP impact PTP impact
Timestamping point Software, so kernel scheduling and interrupt latency add jitter Hardware at the PHY or MAC, so jitter is negligible
Network asymmetry Directly biases the offset estimate — the killer error on WAN paths Partly corrected by delay mechanisms and per-hop correction fields
Residence time in switches Uncorrected; a queued packet is a queued measurement Measured and written into the correction field by transparent clocks
Packet delay variation Averaged out over many samples, slowly Filtered hard; PTP expects mostly-idle or scheduled paths
Oscillator quality Matters between polls Determines holdover when the grandmaster is lost

The practical consequence: NTP degrades gracefully across a congested WAN because it averages many one-second samples. PTP demands a well-behaved path, and its accuracy collapses when packets queue behind bulk traffic. If you cannot keep the PTP path clean, you cannot keep nanosecond accuracy, no matter how good the grandmaster is.

Timestamping and the message exchange

PTP's accuracy comes from measuring four timestamps instead of two. The device records when a Sync leaves the master, when it arrives at the slave, when a Delay_Req leaves the slave and when it arrives back at the master. Offset and path delay fall out of those four numbers.

# Two mechanisms for measuring path delay
# E2E (end to end): slave sends Delay_Req straight to the grandmaster
# P2P (peer to peer): each link measures its own delay with Pdelay_Req

# Two ways of stamping the Sync
# one-step: the timestamp is inserted on the fly at the egress port
# two-step: a Follow_Up message carries the egress timestamp

! Typical switch-side options
set protocols ptp clock-mode boundary
set protocols ptp slave interface ge-0/0/0.0
set protocols ptp master interface ge-0/0/1.0
set protocols ptp transparent-clock
run show ptp clock
run show ptp port
run show ptp statistics

Two-step was historically the only option and is still the safest to interoperate with; one-step reduces message count but requires hardware that can write the timestamp into the frame as it leaves. In a mixed-vendor network, stick with two-step unless you have tested the combination.

PTP profiles matter more than the protocol

IEEE 1588 defines a toolbox; profiles define which tools you use. Two devices can both speak PTP and still never lock, because they are using different profiles with different message rates, transports and optional features.

Profile Domain Used for Typical accuracy
Default (E2E, L2 multicast) 0 Lab and general purpose Sub-microsecond with hardware stamps
G.8275.1 (telecom, full timing support) 24 Mobile backhaul, 5G fronthaul Tens of nanoseconds
G.8275.2 (telecom, partial support) 44 Where every hop cannot be a boundary clock Better than 1 us
C37.238 (power utility) 0 Substation automation 1 us
802.1AS / gPTP (TSN, AV) 0 Audio video bridging, industrial TSN Sub-microsecond
SMPTE 2059 / AES67 0 Broadcast and studio audio Sub-microsecond

Profile mismatch is the second most common "PTP is not locking" cause after a domain mismatch. Check the profile name, the domain number, the transport (L2 versus UDPv4) and the announce interval on both ends before you look at anything else.

NTP is not a lesser protocol, it is a different requirement

For the vast majority of systems — servers, application logs, DNS, Kubernetes, databases, backups — sub-millisecond is not a requirement; consistency is. NTP with a good local reference does that extremely well.

# A solid chrony client configuration
server 10.10.0.10 iburst prefer
server 10.10.0.11 iburst
pool ntp.example.org iburst maxsources 3
driftfile /var/lib/chrony/drift
makestep 1.0 3
rtcsync
logdir /var/log/chrony

chronyc tracking
chronyc sources -v
chronyc sourcestats -v
timedatectl status

Three details separate a working NTP estate from a fragile one. Use iburst so a client reaches useful accuracy in seconds rather than hours. Run at least four upstream sources so the selection algorithm can outvote a falseticker. And treat anycast or a well-monitored local stratum-1 as the reference, not a random public pool — a device that syncs to an arbitrary internet address is one routing change away from losing time.

Also decide your leap-second policy in advance. Most Linux and network estates smear the leap second over hours; mixing smeared and unsmeared sources in the same environment produces a step that breaks log correlation and can confuse applications. Choose one approach, document it, and apply it everywhere.

Security: NTS and rogue grandmasters

Plain NTP is trivially spoofable — an attacker who can answer first can shift your clock, and shifted clocks break TLS validation, Kerberos and audit trails. NTPsec and chrony support Network Time Security (NTS), which authenticates the server and the packets.

# NTS-capable client (chrony 4.x)
server time.cloudflare.com iburst nts
ntsdump -c time.cloudflare.com
chronyc authdata

PTP has its own problem: the best-master clock election is an availability mechanism, not an authentication one. A rogue grandmaster with a better clock class wins the election and feeds the whole domain. The mitigations are operational rather than cryptographic — separate domains for production and lab, announce and sync from a known grandmaster list where the platform supports it, and monitor the parent clock identity so an unexpected change raises an alert before the traffic does.

Monitoring accuracy, not just lock state

"It says locked" is not a measurement. Both protocols expose the numbers that tell you whether the design is actually holding.

# PTP
pmc -u -b 0 'GET TIME_STATUS_NP'
pmc -u -b 0 'GET PARENT_DATA_SET'
phc_ctl /dev/ptp0 get
ethtool -T ens1f0          # confirm hardware timestamping is in use

# NTP
chronyc tracking | grep -E "System time|Last offset|Frequency"
chronyc sourcestats -v
ntpq -p

For PTP, watch the offset, the gmPresent flag and the parent-port identity — the parent changing without a planned change window is the early warning of a rogue or failing grandmaster. For NTP, watch the offset and the frequency correction over time; a slowly growing offset or a drifting frequency estimate points at a failing oscillator, and a growing root dispersion points at a path problem.

Choosing, in one table

Requirement Choose Why
Log correlation, general IT, DNS, K8s NTP Milliseconds are enough and operation is cheap
Financial trading, regulatory timestamping PTP Microsecond-to-nanosecond ordering is mandated
5G / mobile backhaul, TDD radio synchronisation PTP (G.8275.x) Phase alignment, not just frequency
Broadcast, AV over IP, TSN PTP (802.1AS / SMPTE) Media devices expect a common time base
Power substation automation PTP (C37.238) Profile defined for the equipment
Wide area across the internet NTP, or NTP + PTP where supported PTP accuracy is not achievable across unknown paths

The most common expensive mistake is building PTP where NTP would have been sufficient. The second most common is running PTP without a plan for the path — no hardware timestamping, no transparent or boundary clocks, and a congested uplink — and then wondering why accuracy is only "better than NTP, barely". Decide the requirement from the application backwards, and let that decide the protocol rather than the other way round.

Related reading: PTP boundary clock vs transparent clock, chrony time sync and drift troubleshooting, Cisco IOS NTP configuration and fabric-wide PTP with Arista AVD.

原文链接:https://alliedtelesis.com/sites/default/files/documents/configuration-guides/ptp_feature_overview_guide_revb.pdf