PTP (IEEE 1588) vs NTP vs Chrony: Time Synchronization for Industrial IoT and Distributed Systems

PTP (IEEE 1588) vs NTP vs Chrony: Time Synchronization for Industrial IoT and Distributed Systems

PTP vs NTP vs Chrony: Time Synchronization for Industrial IoT and Distributed Systems

Most distributed systems fail on time long before they fail on bandwidth. A historian merges two sensor streams and the cause appears to follow its effect. A motion controller samples a drive 300 microseconds late and a loop that was stable on the bench starts to ring. A fleet of edge gateways stamps events with clocks that disagree by 40 milliseconds, and the root-cause analysis of a plant trip becomes guesswork. The question behind all of these is the same: how closely can independent clocks be made to agree, and what does it cost? The honest answer in the PTP vs NTP debate is that the two protocols solve different problems at different price points, and chrony sits in the middle as the best way to run NTP and, increasingly, to consume PTP hardware clocks.

This article explains the mechanism that every synchronisation protocol shares, then shows where NTP, chrony and IEEE 1588 Precision Time Protocol diverge: timestamping location, network support, topology, profiles and failure modes. It includes runnable configuration for ptp4l, phc2sys and chrony, a synthetic simulation of how jitter and path asymmetry become offset error, and a decision matrix you can apply to a plant or a cloud fleet.

What this covers: the two-way exchange and its asymmetry assumption, NTP stratum and chrony behaviour, PTP messages and clock types, the best master clock algorithm, hardware versus software timestamping, linuxptp and chrony configuration, standard profiles, GNSS holdover and security, cloud time services, failure modes, and a decision matrix.

Context and Background

Every clock-synchronisation protocol in common use estimates the offset between two clocks by exchanging timestamped messages. The mathematics has not changed since David Mills formalised it for the Network Time Protocol: a client records when it sent a request, the server records when it received it and when it replied, and the client records when the reply arrived. Four timestamps give two unknowns, offset and round-trip delay, and one assumption that cannot be measured from the exchange itself: the path is symmetric.

NTP is the oldest protocol still in daily use on the internet. The current specification, RFC 5905 (NTPv4), describes a hierarchy of servers organised by stratum and states that on a fast local network, clients can typically be synchronised to within a few hundred microseconds, while across the public internet the error is more likely tens of milliseconds. Those figures are specification-level typical values, not guarantees. In practice, software timestamping, interrupt latency and asymmetric routing set the floor.

IEEE 1588, the Precision Time Protocol, was created for test-and-measurement and industrial automation, where microseconds and below mattered. It was first published in 2002, revised in 2008 as version 2 (the one that most deployed equipment speaks), and revised again in 2019 as version 2.1. The key difference is not a cleverer formula. PTP moves timestamping into the network interface hardware and, in the best deployments, makes every switch on the path participate in the protocol so that queueing delay is removed or measured instead of guessed.

Chrony is a modern NTP implementation for Linux and other Unix-like systems that also happens to be the practical glue between the two worlds. The chrony project documents that it synchronises faster and more accurately than the classic ntpd in most conditions, handles intermittent connectivity better, never steps the clock in its default configuration, supports Network Time Security (RFC 8915), and supports hardware timestamping on Linux for both received and transmitted NTP packets. It can also read a PTP hardware clock as a reference, which is how many production systems combine PTP-grade local time with ordinary applications.

Time synchronisation is a prerequisite for the deterministic networks covered elsewhere on this site. If you are designing converged control networks, read the time-sensitive networking reference architecture alongside this article: TSN schedules are meaningless without a shared timebase, and IEEE 802.1AS, discussed below, is a PTP profile. The primary sources for the claims here are RFC 5905, the chrony project documentation, the linuxptp documentation and the IEEE 1588 profile list, all linked in Further Reading.

How Clock Synchronisation Works: Offset, Delay and the Symmetry Assumption

Direct answer: Two clocks are synchronised by exchanging four timestamps. The offset is half the sum of the forward and backward timestamp differences, and the round-trip delay is the total elapsed time minus the remote processing time. The offset estimate is exact only if the forward and return path delays are equal, so any asymmetry or timestamp jitter becomes error.

PTP vs NTP shared two-way exchange: four timestamps T1 to T4 yield offset and delay under a symmetric path assumption

Figure 1: The two-way exchange used by NTP and by PTP’s end-to-end delay mechanism. Offset is half of (T2 minus T1) plus (T3 minus T4).

Figure 1 shows the exchange. RFC 5905 defines the client-side timestamps T1 (request departs) and T4 (reply arrives) and the server-side timestamps T2 (request arrives) and T3 (reply departs). The offset estimate is theta = 1/2 [(T2 – T1) + (T3 – T4)] and the round-trip delay is delta = (T4 – T1) – (T3 – T2). PTP’s end-to-end mechanism uses the same four times, with different names: Sync and Delay_Req carry the same information.

Why symmetry is the invisible assumption

Suppose the true one-way delay from client to server is d1 and from server to client is d2, and the clocks differ by an unknown offset. The algebra gives an estimated offset equal to the true offset plus (d1 – d2)/2. Nothing in the four timestamps lets you separate the two delays, so the protocol assumes d1 equals d2. If a path is asymmetric by 40 microseconds, every measurement is wrong by 20 microseconds, no matter how many samples you average. Averaging reduces random noise; it cannot remove a bias.

This is why a one-microsecond clock in a data centre and a one-microsecond clock in a plant need different engineering. Random jitter from interrupts, queues and cache effects can be filtered. Systematic asymmetry from cable length differences, fibre wavelength dispersion on single-strand links, or a switch that forwards faster in one direction has to be eliminated at the physical layer or calibrated and compensated explicitly. Propagation in copper or fibre is roughly 5 nanoseconds per metre, so a 10 metre difference in transmit and receive path lengths alone is about 50 nanoseconds of asymmetry and 25 nanoseconds of offset error.

Where timestamp error comes from

The error budget of any protocol is the sum of several terms. The timestamp point error is the difference between the instant a packet crosses the wire and the instant software reads a clock. Queueing delay in switches and routers varies with load. Path asymmetry is the bias term above. Oscillator wander is how far the local clock drifts between corrections, which depends on the quality of the crystal, its temperature environment and the poll interval. Finally the servo itself, the control loop that decides how to slew the clock, adds its own noise or lag.

A common crystal oscillator in a server may be off by tens of parts per million (ppm) uncorrected. One ppm is one microsecond of drift per second. A temperature-compensated crystal oscillator (TCXO) is typically better, and an oven-controlled one (OCXO) better again, with figures measured in parts per billion (ppb). One ppb is one microsecond of drift per 1,000 seconds. These orders of magnitude explain why PTP deployments send Sync messages many times per second and why NTP, polling at intervals of 64 seconds to 1,024 seconds, relies on very good frequency estimation instead of fast correction.

Offset versus frequency

A clock has a phase (what time is it) and a frequency (how fast does it run). The protocols synchronise phase, but a good servo controls frequency, because correcting rate keeps phase small between messages. Chrony and ptp4l both estimate frequency error and adjust the clock rate using the kernel’s adjtimex facility or the hardware clock’s own frequency adjustment. Stepping the time (a discontinuous jump) is reserved for the initial correction, because jumps break monotonic assumptions in databases, schedulers and control code. Chrony’s default behaviour is to slew only; the makestep directive explicitly allows a step in the first few updates after start.

NTP and Chrony in Depth

NTP’s design goals are scale and robustness over the internet. A stratum-1 server is directly connected to a reference clock such as a GNSS receiver; stratum 2 syncs to stratum 1, and so on. RFC 5905 defines stratum 1 as primary, 2 through 15 as secondary, 16 as unsynchronised, and reserves the rest. Stratum is a distance metric, not a quality metric: a stratum-2 server on the same LAN with a stable path can be better than a distant stratum-1.

What the NTP algorithms do

A client polls several servers. Per-server filters pick the best recent samples, which tends to be the one with the lowest round-trip delay since queueing only adds delay. A selection algorithm discards falsetickers by finding the largest group of servers whose confidence intervals overlap, then a combine step weights survivors. A clock-discipline loop then adjusts frequency and phase. This pipeline is robust to a few lying servers and to transient congestion, which is exactly what you want across a wide area network and exactly why NTP does not reach PTP-class accuracy: it has no way to remove variable queueing from the sample, only to filter it.

What chrony changes in practice

The chrony comparison page highlights several operational differences that matter more than raw accuracy numbers. Chrony copes with intermittent reachability and with hosts that hibernate or run on unstable clocks, such as virtual machines, because it can slew over a wide frequency range and models oscillator behaviour quickly. It has a smaller footprint. It supports kernel and hardware timestamping on Linux, which classic ntpd does not, and Network Time Security. It lacks the broadcast, multicast and manycast modes and the many hardware reference-clock drivers of the reference ntpd; chrony gets reference time through shared memory (SHM), a socket (SOCK), PPS, RTC or PTP hardware clock (PHC) interfaces, typically fed by gpsd or by a PHC.

The most important chrony feature for industrial and low-latency work is the combination of NIC hardware timestamping and a local NTP server. With hardware timestamps on a quiet LAN, chrony can often hold a host within a few microseconds of a nearby server, though the figure depends on the NIC, switch behaviour and load and should be measured, not assumed. That is a very different regime from the tens of milliseconds expected across the internet.

Leap seconds and smearing

UTC occasionally inserts a leap second, and different systems handle it differently. Chrony can step, slew, ignore, or let the kernel handle it, and as a server can smear the correction over a long window. Google and AWS public NTP services smear instead of inserting a visible 23:59:60. AWS documents that its local and public NTP endpoints smear but its PTP hardware clock does not, and recommends against mixing smeared and non-smeared sources in a single client during a leap event. That sentence alone is the source of many outages waiting to happen, and we return to it in the failure-modes section.

PTP carries a different timescale. Its epoch is the PTP epoch tied to TAI, with a currentUtcOffset field that tells receivers how many seconds TAI is ahead of UTC (37 seconds at the time of writing; verify the current value against IERS bulletins). Converting PTP time to UTC for system use is therefore an explicit step, which ptp4l and phc2sys manage. Leap seconds appear as announcements in PTP Announce messages rather than as a smear.

PTP Architecture: Messages, Clock Types and the Best Master Clock Algorithm

Direct answer: PTP synchronises clocks in a domain through a grandmaster clock chosen by the best master clock algorithm (BMCA). Boundary clocks terminate and regenerate PTP in each switch, transparent clocks add the residence time to a correction field, and ordinary clocks sit at the leaves. Hardware timestamping at the port removes most software jitter.

IEEE 1588 boundary clock and transparent clock topology between a GNSS grandmaster and ordinary clocks on industrial IoT nodes

Figure 2: A PTP domain with a GNSS-fed grandmaster, a boundary clock, a transparent clock, ordinary clocks and a backup grandmaster that takes over through the BMCA.

The message set

PTP defines a small set of message types. Announce messages advertise a clock’s properties so the BMCA can pick the best source. Sync messages carry (or, in two-step operation, trigger the later reporting of) the time the master sent them. In two-step mode the precise transmit timestamp arrives in a Follow_Up message, because the hardware only knows the exact egress time after the frame has left the port. Delay_Req and Delay_Resp implement the end-to-end delay mechanism: the slave timestamps its request and the master reports when it arrived. Pdelay_Req, Pdelay_Resp and Pdelay_Resp_Follow_Up implement the peer-to-peer mechanism, which measures the link delay hop by hop instead of the whole path. Management and Signaling messages configure and negotiate.

Event messages (Sync, Delay_Req and the Pdelay messages) are the ones timestamped in hardware; general messages (Announce, Follow_Up, Delay_Resp, Management, Signaling) are not time-critical. Over UDP the event messages use port 319 and the general messages port 320, while Layer 2 transport uses EtherType 0x88F7. The L2 option matters in industrial networks because it removes IP and UDP header processing and avoids routers entirely.

A one-step clock inserts the transmit timestamp into the Sync message as it leaves the port, so no Follow_Up is needed. Two-step is easier to build and is more common. The linuxptp documentation lists time_stamping modes including hardware, software, legacy, onestep and p2p1step; the default is hardware.

Ordinary, boundary and transparent clocks

An ordinary clock has a single PTP port and is either a grandmaster (source) or a slave (consumer). A boundary clock has several ports; it synchronises to the grandmaster on one port, then acts as a master on the others, regenerating timing from its own disciplined clock. This isolates each segment: queueing in an upstream switch cannot propagate past a boundary clock, and the number of slaves a grandmaster must serve is bounded. The cost is that each boundary clock adds its own servo noise, so long chains accumulate error and are limited by profile (the ITU-T telecom profiles specify budgets per hop).

A transparent clock does not synchronise to anything. It measures how long each PTP event message spends inside the device, the residence time, and adds that to a correction field in the message, so the slave can subtract it. End-to-end transparent clocks correct only residence time; peer-to-peer transparent clocks also measure and correct the link delay of the incoming port, which makes topology changes faster and tolerates asymmetric-but-measured links. Transparent clocks keep a pure forwarding behaviour and scale well, but they require every switch in the path to be PTP-aware; a single non-aware switch that queues PTP packets with ordinary traffic can undo the whole benefit.

The best master clock algorithm

Every clock announces its attributes and each port runs the BMCA over the Announce messages it receives, so the network converges on a single grandmaster per domain without a configuration server. The comparison order is: priority1, clockClass, clockAccuracy, offsetScaledLogVariance, priority2 and finally clockIdentity as a tie-break. Lower is better at each step. The linuxptp defaults are priority1 of 128 and a clockClass of 248, meaning a clock that is not traceable to a primary reference, so an operator who wants a particular device to win sets priority1 low, and a device fed by GNSS advertises a lower clockClass (6 is the value for a clock synchronised to a primary reference) while it is locked.

Two operational consequences matter. First, the BMCA is automatic, so a rogue device announcing priority1 0 will become grandmaster; the protocol has no built-in authentication, which is why access control and port-level filtering matter. Second, when the grandmaster loses GNSS, it degrades its clockClass to indicate holdover, and a backup grandmaster with better class can take over, but the switchover takes several announce intervals. The announceReceiptTimeout default is three missed announces, so with one Announce per second the takeover is measured in seconds.

PTP time, UTC and the system clock

On a Linux host a PTP-capable NIC exposes a PTP hardware clock, a free-running counter in the NIC that appears as /dev/ptpN. ptp4l disciplines that PHC to the grandmaster using hardware timestamps. The operating system clock (CLOCK_REALTIME) is a separate clock; it is disciplined from the PHC by phc2sys, which also applies the UTC to TAI offset. The linuxptp documentation notes that with hardware timestamping, ptp4l relies on phc2sys to maintain the correct offset between UTC and PTP time. This two-stage structure, NIC clock first and system clock second, is the central implementation detail of PTP on Linux and is a frequent source of confusion: ptp4l reporting a sub-microsecond offset says nothing about the offset of the application’s clock.

Software vs Hardware Timestamping and the PTP Hardware Clock

Timestamping location is the single largest factor in the PTP vs NTP accuracy gap. In a software-timestamped exchange, the packet arrives at the NIC, is DMA-ed into memory, raises an interrupt, travels through the kernel network stack and is stamped when software reads the system clock. Each step has variable latency from cache state, interrupt coalescing, scheduler pressure and competing traffic. Tens of microseconds of jitter on a busy host is ordinary.

Hardware timestamping vs software timestamping paths from the wire to the clock reading, showing where jitter enters

Figure 3: Software timestamps are taken after the interrupt and kernel path; hardware timestamps are taken at the port and read from the NIC’s PHC counter.

With hardware timestamping, the MAC or PHY latches the NIC’s own counter at the moment a recognised frame crosses the port, and the stamp is handed to software later through the socket error queue for transmit and as ancillary data for receive. Linux exposes this through the SO_TIMESTAMPING socket option and the ethtool -T command, which lists supported capabilities. The timestamp error reduces to the counter resolution (nanoseconds, determined by the clock frequency of the NIC), residual PHY latency variation and any fixed offset between the stamping point and the true reference plane, which can be calibrated per device.

The cost of getting this wrong

NTP can use hardware timestamping too; chrony does with the hwtimestamp directive, and a chrony-to-chrony LAN with hardware stamps behaves much more like PTP than like internet NTP. The real difference then becomes the network: NTP has no equivalent of boundary or transparent clocks, so switch queueing remains in the sample. On an idle network that is a few microseconds; on a congested one it is hundreds of microseconds or worse. PTP with transparent or boundary clocks is designed to remove exactly that term.

Quantifying the effect with a synthetic model

The following script (also saved as assets/sim_two_way.py) is a Monte Carlo model of the four-timestamp exchange. Every number it prints is synthetic: delays and jitter are inputs I chose to illustrate the mechanism, not measurements from any device. Each one-way delay is a fixed propagation time plus exponentially distributed queueing jitter.

"""Synthetic two-way time transfer: how jitter and path asymmetry reach the offset estimate."""
import random
import statistics as st

TRUE_OFFSET = 250e-6          # client clock is 250 microseconds ahead of the server
TURNAROUND = 20e-6            # server processing time between T2 and T3

def estimate(fwd, rev, jitter, rng):
    d1 = fwd + rng.expovariate(1 / jitter)   # client -> server one-way delay
    d2 = rev + rng.expovariate(1 / jitter)   # server -> client one-way delay
    t1 = TRUE_OFFSET                          # client clock at send
    t2 = d1                                   # server clock at receive
    t3 = t2 + TURNAROUND                      # server clock at reply
    t4 = t3 + d2 + TRUE_OFFSET                # client clock at arrival
    theta = 0.5 * ((t2 - t1) + (t3 - t4))     # RFC 5905: server minus client
    delay = (t4 - t1) - (t3 - t2)
    return -theta, delay                      # report client minus server

def run(name, fwd, rev, jitter, n=5000, seed=7):
    rng = random.Random(seed)
    errs = []
    for _ in range(n):
        off, _ = estimate(fwd, rev, jitter, rng)
        errs.append(off - TRUE_OFFSET)
    print(f"{name:32s} mean error {st.mean(errs)*1e6:8.2f} us   stdev {st.pstdev(errs)*1e6:7.2f} us")

run("symmetric, jitter 50 us", 100e-6, 100e-6, 50e-6)
run("symmetric, jitter 0.1 us", 100e-6, 100e-6, 0.1e-6)
run("asymmetric 40 us, jitter 0.1 us", 120e-6, 80e-6, 0.1e-6)
run("asymmetric 400 ns, jitter 0.1 us", 100.2e-6, 99.8e-6, 0.1e-6)

Running it with seed 7 gave this output on my sandbox (illustrative only):

symmetric, jitter 50 us          mean error    -0.04 us   stdev   35.18 us
symmetric, jitter 0.1 us         mean error    -0.00 us   stdev    0.07 us
asymmetric 40 us, jitter 0.1 us  mean error   -20.00 us   stdev    0.07 us
asymmetric 400 ns, jitter 0.1 us mean error    -0.20 us   stdev    0.07 us

Three lessons fall out. With large symmetric jitter, individual samples are poor (35 microsecond standard deviation) but the mean is unbiased, so a filter with enough samples recovers the offset; this is the NTP regime. With tiny jitter the per-sample error collapses to 0.07 microseconds; this is the PTP-with-hardware-stamps regime. And with a deterministic asymmetry, the error is exactly half the asymmetry (20 microseconds for 40; 0.2 for 0.4) and no amount of sampling helps. Real systems have all three effects at once, but the order of magnitude of each term tells you where to spend engineering effort.

Running It: linuxptp, chrony and Monitoring

The following configurations are minimal, realistic starting points. Interface names, domain numbers and thresholds must be adapted, and the exact option names should be checked against the man pages of the versions you deploy, since linuxptp has renamed some options over time (for example, clientOnly replacing the older slaveOnly).

A PTP slave with ptp4l and phc2sys

First confirm that the NIC supports hardware timestamping and has a PHC:

ethtool -T eth0
# Look for: hardware-transmit, hardware-receive, hardware-raw-clock
# and "PTP Hardware Clock: <n>" (n >= 0)

A client-only configuration for a Layer 2 transport with end-to-end delay measurement, as used by several industrial and telecom profiles:

# /etc/linuxptp/ptp4l-client.cfg
[global]
domainNumber            24
clientOnly              1
time_stamping           hardware
network_transport       L2
delay_mechanism         E2E
logAnnounceInterval     0
announceReceiptTimeout  3
logSyncInterval         -4      # 16 Sync per second
logMinDelayReqInterval  -4
summary_interval        4       # print statistics every 16 seconds
tx_timestamp_timeout    20
step_threshold          0.00002
first_step_threshold    0.00002

[eth0]

Run ptp4l and then phc2sys, which waits for ptp4l (the -w flag) and copies the PHC time to the system clock with the right UTC offset:

sudo ptp4l -f /etc/linuxptp/ptp4l-client.cfg -m
sudo phc2sys -s eth0 -c CLOCK_REALTIME -w -m

In the ptp4l log the servo state is shown as s0 (unlocked), s1 (stepping the clock) and s2 (locked). Reaching s2 and holding it is the first acceptance criterion. Query the daemon through the management interface with the pmc tool:

sudo pmc -u -b 0 'GET TIME_STATUS_NP'     # offset from master, gmPresent, gmIdentity
sudo pmc -u -b 0 'GET PORT_DATA_SET'      # port state: MASTER, SLAVE, LISTENING
sudo pmc -u -b 0 'GET PARENT_DATA_SET'    # grandmaster identity and variance

A grandmaster fed by a GNSS receiver is similar but sets priority1 low, specifies a clockClass appropriate to the time source, and typically uses ts2phc to align the NIC’s PHC to the receiver’s pulse-per-second output. A pure software grandmaster running on a server is acceptable for a lab; it is not a substitute for a hardware grandmaster with a disciplined oscillator in a plant that depends on holdover.

Chrony as an NTP server and a PHC consumer

On servers that need accurate system time without running a full PTP stack, chrony with hardware timestamping against local servers is often enough. And on a host already running ptp4l, chrony can use the PHC as a refclock instead of phc2sys, which gives chrony’s filtering and fallback behaviour:

# /etc/chrony/chrony.conf  (illustrative production starting point)

# Option A: NTP from local stratum-1 appliances, with hardware timestamps and NTS to a public pool as a sanity check
server ntp1.plant.example iburst minpoll 2 maxpoll 4 xleave
server ntp2.plant.example iburst minpoll 2 maxpoll 4 xleave
server time.cloudflare.com iburst nts
hwtimestamp eth0
rtcsync
makestep 1.0 3                  # allow a step only in the first 3 updates
driftfile /var/lib/chrony/drift
leapsecmode slew
maxslewrate 1000                # ppm

# Option B: use the PTP hardware clock as a reference (ptp4l disciplines the PHC)
# refclock PHC /dev/ptp0 poll 0 dpoll -2 offset 0 tai prefer trust

# Serve time to the local control network
allow 10.20.0.0/16
local stratum 10 orphan         # keep serving a consistent time if all sources are lost

The tai option on a PHC refclock tells chrony that the clock runs on TAI so it applies the leap-second offset; the offset 0 may need adjusting if the PHC and system time have a known fixed offset. The interleaved mode (xleave) lets a chrony server and client use the transmit timestamp of the previous packet, which is how hardware transmit timestamps become useful in NTP without PTP’s two-step Follow_Up message. Verify behaviour with:

chronyc tracking          # System time offset, RMS offset, frequency, skew, root delay
chronyc sources -v        # Per-source reachability, last sample, and selection state
chronyc sourcestats       # Per-source frequency and offset statistics
chronyc ntpdata           # Per-association timestamping mode, delays and jitter

A monitoring script you can alert on

Offset alone is not enough; you also need to know whether the clock is synchronised, who the grandmaster is, and whether the offset is stable. This small Python script reads chrony’s CSV output and flags thresholds. Adapt the thresholds to your tolerance.

#!/usr/bin/env python3
"""Emit a chrony health summary suitable for a Prometheus textfile collector."""
import subprocess
import sys

OFFSET_WARN_US = 50.0
STRATUM_MAX = 4

def tracking():
    out = subprocess.check_output(["chronyc", "-c", "tracking"], text=True).strip()
    f = out.split(",")
    # Fields per chrony docs: refid, refname, stratum, ref_time, system_time_offset,
    # last_offset, rms_offset, freq_ppm, resid_freq_ppm, skew_ppm, root_delay,
    # root_dispersion, update_interval, leap_status
    return {
        "refname": f[1],
        "stratum": int(f[2]),
        "system_offset_s": float(f[4]),
        "rms_offset_s": float(f[6]),
        "freq_ppm": float(f[7]),
        "root_delay_s": float(f[10]),
        "root_dispersion_s": float(f[11]),
        "leap": f[13],
    }

def main():
    t = tracking()
    offset_us = abs(t["system_offset_s"]) * 1e6
    status = 0
    if t["leap"] == "Not synchronised" or t["stratum"] > STRATUM_MAX:
        status = 2
    elif offset_us > OFFSET_WARN_US:
        status = 1
    print(f"chrony_offset_us {offset_us:.3f}")
    print(f"chrony_rms_offset_us {t['rms_offset_s']*1e6:.3f}")
    print(f"chrony_frequency_ppm {t['freq_ppm']:.3f}")
    print(f"chrony_root_distance_us {(t['root_delay_s']/2 + t['root_dispersion_s'])*1e6:.1f}")
    print(f"chrony_status {status}")
    return status

if __name__ == "__main__":
    sys.exit(main())

The root distance printed there (half the root delay plus root dispersion) is chrony’s upper bound on the error of the system clock relative to the primary reference, assuming the path asymmetry stays within the delay. It is a more honest health metric than the instantaneous offset, because a low offset with a large root distance means the number may be luck.

Profiles, Holdover, Security and Cloud Time

A profile is a set of selections from the IEEE 1588 option space (transport, delay mechanism, message rates, domain numbers, BMCA parameters, accuracy requirements) that makes equipment from different vendors interoperate for a given industry. The IEEE 1588 working group maintains a list of profiles that includes the generic default profiles, ITU-T G.8265.1 (frequency), G.8275.1 and G.8275.2 (phase and time), the IETF enterprise profile draft, SMPTE ST 2059-2, AES67, IEEE 802.1AS-2020, IEEE C37.238-2017, IEC/IEEE 61850-9-3 and several IEC 62439-3 annexes for automation networks.

The profiles that matter for industrial IoT

Profile Domain Transport and delay mechanism Typical use Notes
IEEE 802.1AS-2020 (gPTP) Industrial, automotive, AV Layer 2, peer-to-peer, two-step TSN time base, in-vehicle and factory Ethernet Every hop must be gPTP-capable; basis for TSN scheduling
IEC/IEEE 61850-9-3 and IEEE C37.238 Power utilities Layer 2, peer-to-peer Substation process bus, synchrophasors Utility-specific profile; check the standard text for its time-error requirements
ITU-T G.8275.1 Telecom Layer 2 multicast, end-to-end, boundary clock at every node 5G fronthaul, phase and time Requires full on-path support; domain 24 by default
ITU-T G.8275.2 Telecom UDP/IP unicast Partial on-path support Tolerates non-PTP-aware hops with worse accuracy
SMPTE ST 2059-2 Broadcast UDP/IP Studio media over IP Used with ST 2110
Default profiles (1588-2019) General UDP/IP, E2E or P2P Labs, generic networks Rarely enough for interoperability on its own

For plant networks the one to know is IEEE 802.1AS. It is a restricted, strictly Layer 2 profile with peer-to-peer delay measurement on every link, intended to provide the common timebase that IEEE 802.1Qbv time-aware scheduling needs. If you are building a converged deterministic network, see how the timebase interacts with scheduled traffic in our OPC UA over TSN reference architecture and, for protocol-level contrasts between fieldbuses, PROFINET vs EtherCAT vs OPC UA FX. The IRT variant of PROFINET carries its own synchronisation mechanism, compared with TSN in the PROFINET IRT and TSN guide.

The telecom profile G.8275.1 illustrates how strict profiles are. One independent profile summary describes it as Layer 2 only with multicast MAC addresses, end-to-end delay measurement, 16 Sync and 16 Delay_Req per second, one Announce per second and domain 24 as the recommended default, with every intermediate device a boundary clock or a suitable transparent clock. The same summary cites a plus or minus 1.5 microsecond time-error budget for 5G fronthaul. Treat these as a profile example, and consult the ITU-T recommendation text before relying on any number.

GNSS, grandmasters and holdover

Most grandmasters derive time from GNSS (GPS, Galileo and others). A GNSS receiver delivers a 1 pulse-per-second (PPS) signal accurate to tens of nanoseconds under good conditions, plus a time-of-day message. The grandmaster disciplines a local oscillator to the PPS, and that oscillator provides holdover when the sky view is lost, or the signal is jammed or spoofed. Holdover quality depends entirely on the oscillator: with a frequency error of 1 ppb, the phase error grows by 1 microsecond per 1,000 seconds; at 10 ppb it grows by 1 microsecond in 100 seconds. Those are illustrative arithmetic, not a specification of any product. A rubidium or good OCXO holdover is therefore the difference between riding out a 15 minute outage and a plant-wide timing fault.

GNSS is also a vulnerability. Jamming is simple and spoofing is feasible, and a plant that trusts one rooftop antenna has a single point of failure with a large blast radius. Mitigations include a multi-constellation receiver with anti-spoofing features, redundant grandmasters from independent antennas, a terrestrial reference such as a time service delivered over fibre, and monitoring that cross-checks GNSS time against a second source and alarms on disagreement. A PTP network where the BMCA is allowed to follow the grandmaster’s degrading clockClass will fall back to a backup automatically; check that the backup is actually better than holdover before relying on it.

Security: NTS for NTP, and the PTP gap

NTP security is further along. Network Time Security (RFC 8915) protects client-server NTP: an initial TLS 1.3 handshake on TCP port 4460 (NTS-KE) establishes keys and issues cookies, and subsequent NTP packets carry extension fields (unique identifier, cookie, and an authenticator using AEAD) that authenticate responses without per-client server state. Chrony and NTPsec support it; the classic ntpd does not, and the older symmetric-key and Autokey mechanisms are weaker or obsolete. NTS authenticates a server’s response but cannot stop a delay attack: an on-path attacker can still delay packets and bias the offset, because delaying is indistinguishable from a long path.

PTP has traditionally had weak protection in practice. IEEE 1588-2019 contains a security annex describing an experimental authentication TLV approach, and the IETF has drafts on using NTS key establishment for PTP. At the time of writing I could not confirm that any of this is widely deployed in plant equipment, so treat PTP security as unresolved and defend the network instead: restrict who can send Announce messages (port-level filtering, VLAN isolation of timing traffic), use priority and clockClass settings with known grandmaster identities, and monitor the grandmaster identity and offset so that an unexpected change raises an alarm. Defence by network design is the realistic baseline for the foreseeable future.

Cloud time services

On public cloud, the practical question is which source to point chrony at. AWS operates the Amazon Time Sync Service, backed by a fleet of satellite-connected and atomic reference clocks in each Region. Instances reach it over a local NTP endpoint and over public pools, and on supported instances through a PTP hardware clock with 64-bit nanosecond-precision hardware timestamps added to incoming packets. AWS’s documentation states that the NTP endpoints smear leap seconds while the PTP hardware clock does not, and recommends against configuring both smeared and non-smeared sources in the same client during a leap event. Google Cloud and Microsoft Azure offer their own time sources; check their current documentation for addresses, smearing policy and PTP support, which I have not re-verified for this article. A cloud-hosted twin or analytics platform generally wants chrony pointed at the provider’s local source, because the path is short and the source is smeared consistently.

PTP vs NTP vs Chrony: Choosing the Right Tool

Direct answer: Use chrony with NTP when you need millisecond to tens of microsecond accuracy across ordinary networks or the internet. Use PTP with hardware timestamping and PTP-aware switches when you need microsecond to sub-microsecond accuracy inside a controlled network. Use both: PTP for the local timebase, chrony or phc2sys to deliver it to applications.

PTP vs NTP decision flow from required accuracy and network support to chrony, NTS or PTP with hardware timestamping

Figure 4: A decision flow for time synchronisation: required accuracy and on-path PTP support determine whether chrony, a cloud time service, or PTP is the right tool.

Decision matrix

Criterion NTP (classic ntpd) Chrony PTP (linuxptp)
Specification typical accuracy Hundreds of microseconds on a fast LAN, tens of ms over the internet (RFC 5905) Similar over WAN; microseconds with hardware timestamping on a quiet LAN (measure it) Sub-microsecond to nanoseconds with hardware stamps and on-path support; far worse without
Timestamping Software Software or hardware (Linux) Hardware strongly preferred
Network support needed None None Boundary or transparent clocks for best results
Topology Server hierarchy with strata Same, plus local refclock inputs Grandmaster chosen by BMCA, per domain
Security Symmetric keys, Autokey obsolete NTS (RFC 8915) Limited; mostly network-level controls
Operational burden Low Low Higher: profiles, switches, GNSS, monitoring
Typical sweet spot Legacy and appliances General servers, VMs, edge, cloud Plants with TSN or fieldbus, substations, broadcast, 5G
Failure signature Gradual drift, wrong falsetickers Large root distance, source selection flapping Asymmetry bias, grandmaster flips, servo unlock

Rules of thumb that survive contact with reality

Start from the requirement, not the protocol. If your worst-case acceptable skew is 1 millisecond, NTP via chrony is the correct, cheapest answer, and PTP would add switches, GNSS and operational risk for no benefit. If a control loop at 1 kHz needs the controller and drive clocks within tens of microseconds, hardware-timestamped PTP over a PTP-aware network is the credible option. Between those, hardware-timestamped chrony on a well-engineered LAN often gets you most of the way.

Match the whole path, not just the endpoints. A 100 nanosecond grandmaster behind a switch that treats PTP as best-effort traffic delivers a 100 microsecond timebase. Check that every switch is boundary-clock, transparent-clock or gPTP capable, that PTP traffic is not shaped behind bulk flows, and that a firmware update has not silently changed this.

Separate the system clock from the PHC. In most Linux deployments the PHC is disciplined by ptp4l and the system clock by phc2sys or by chrony reading the PHC; running both phc2sys and chrony against the same system clock leads to two controllers fighting. Pick one writer to CLOCK_REALTIME.

Trade-offs, Gotchas, and What Goes Wrong

Asymmetry you cannot see. The estimate trusts equal forward and return delays. Single-fibre bidirectional links use different wavelengths in each direction, which have different group delay; a few hundred metres of fibre can add tens of nanoseconds of asymmetry. Copper cable pair lengths differ, and so do FPGA and PHY paths. Calibrate with a reference, set the delayAsymmetry value where the profile supports it, and re-verify after hardware changes.

Switches that do not participate. A non-PTP-aware switch in the path adds variable queueing, which shows up as large and load-dependent offset error. The failure is intermittent, which makes it nasty: the system looks synchronised at night and drifts at shift change. The ptp4l path-delay figure in the summary output is a quick indicator, since a path delay that varies from sample to sample means queueing is leaking in.

Competing synchronisation loops. Running chrony and phc2sys against CLOCK_REALTIME, or ntpd and chrony together, or a hypervisor time-sync agent alongside a guest NTP client, produces fights that look like random steps. Decide one authority per clock.

Grandmaster flapping and rogue masters. Announce interval and timeout settings that are too aggressive cause oscillation when a link glitches. A misconfigured or hostile device with priority1 set low takes over the domain and the whole plant jumps. Use consistent priorities, lock down who can announce, and alarm on grandmaster identity changes.

Leap second and timescale mixing. Mixing smeared NTP time with unsmeared PTP time (as AWS cautions) gives a half-second discrepancy over a smear window or a one-second discontinuity. Mixing the TAI-based PTP epoch and UTC without applying the offset gives a 37 second error, a classic bug when a PHC reading is used directly as a log timestamp.

Virtualisation. A VM’s clock reads through a hypervisor, and hardware timestamps do not pass through virtual NICs without SR-IOV or paravirtual PHC devices. Pause and resume events and live migration can jump time. In VMs prefer a hypervisor-provided PHC or the cloud provider’s time source and let chrony absorb jitter.

Over-claiming accuracy. Vendors quote best-case figures. An acceptance test should measure offset against an independent reference, such as a PPS output from each device compared on an oscilloscope or time-interval counter, under load, during failover and across temperature. Without that measurement, the number in a datasheet is a hypothesis.

Practical Recommendations

For most enterprise and cloud fleets, deploy chrony everywhere, point it at at least three or four independent sources (local appliances plus provider time services), enable NTS where the sources support it, and never mix smeared and unsmeared sources. Monitor root distance, not just offset.

For plants, start with the requirement table from the control or measurement engineers: which devices need what skew, relative to what reference, and under which failure conditions. Choose a profile that matches the equipment (802.1AS if TSN, the power profile in substations, a telecom profile for 5G fronthaul), then design the network for it: PTP-aware switches, redundant grandmasters on independent antennas, short boundary-clock chains, isolated VLANs for timing traffic and a holdover specification that covers your worst credible GNSS outage.

On Linux endpoints, use hardware timestamping and a PHC, run ptp4l for the NIC clock, and let exactly one process write CLOCK_REALTIME. Record the servo state, offset, path delay and grandmaster identity in your observability stack, and alert on state changes rather than only threshold crossings. Stamp data with PTP/TAI or UTC explicitly and say which in the schema.

  • Define the skew budget per consumer (historian, controller, camera, meter) in microseconds.
  • Verify end-to-end PTP awareness in every switch and the exact firmware versions.
  • Calibrate path asymmetry for fibre and cable lengths; document it.
  • Provision at least two grandmasters on independent GNSS paths; test failover.
  • Choose a single writer for CLOCK_REALTIME on every host.
  • Monitor offset, path delay, root distance, servo state and grandmaster identity.
  • Plan for leap seconds and test smeared vs unsmeared behaviour in staging.
  • Measure against an independent reference at commissioning and after changes.

Frequently Asked Questions

What is the difference between PTP and NTP?

NTP is designed for scale and resilience across arbitrary networks and normally uses software timestamps, giving typical accuracy of hundreds of microseconds on a fast LAN and tens of milliseconds over the internet. PTP (IEEE 1588) uses hardware timestamping and, ideally, PTP-aware switches acting as boundary or transparent clocks, which removes most queueing and software jitter and reaches sub-microsecond accuracy in well-engineered networks.

Is chrony better than ntpd?

For most Linux systems, yes. The chrony project documents faster and more accurate synchronisation, better behaviour with intermittent connections and unstable clocks such as virtual machines, a smaller footprint, hardware timestamping and NTS support. ntpd supports more operating systems, more reference-clock drivers and broadcast, multicast and manycast modes, so legacy environments may still need it. Chrony is the sensible default for new deployments.

Can chrony use hardware timestamping instead of PTP?

Yes. On Linux, chrony supports hardware timestamping of received and transmitted NTP packets with the hwtimestamp directive when the NIC and driver support it, and the xleave option makes transmit timestamps usable. Accuracy on a quiet local network can be a few microseconds, but switch queueing still enters the measurement because NTP has no boundary or transparent clock equivalent. Measure your own result against an independent reference.

What is a boundary clock versus a transparent clock?

A boundary clock terminates PTP on one port, synchronises its own clock to the master, and acts as a new master on its other ports, isolating each segment from upstream queueing. A transparent clock does not synchronise; it measures how long each PTP message spends inside the device and adds that residence time to a correction field so the slave can compensate. Both need hardware support.

Do I need GNSS for PTP?

Not for relative synchronisation: devices in a domain can agree with each other using an internal grandmaster oscillator. You need GNSS or another external reference when the domain must also be traceable to UTC, for example for timestamps shared across sites, regulatory records or telecom phase requirements. A GNSS-fed grandmaster should have a good oscillator for holdover and a second independent source or grandmaster for resilience.

Which time sync should I use in the cloud?

Use chrony pointed at the provider’s local time source, because the path is short and the source is managed. AWS documents a local NTP endpoint and a PTP hardware clock on supported instances, with the NTP endpoints smearing leap seconds and the PTP clock not. Avoid mixing smeared and unsmeared sources in one client, and check the current documentation of your provider before configuring.

Further Reading

Internal:

External primary sources (all fetched and checked while writing):

By Riju — about

Comments

No comments yet. Why don’t you start the discussion?

Leave a Reply

Your email address will not be published. Required fields are marked *