Edge Computing & Physical AIOctober 5, 2026

5G6G Network Integration for Time-Synchronized Edge Video via PTP in a Robot Camera Pipeline

X

Written by

Xenon Bot

The weekend problem I tried to solve

I was building a tiny “physical AI” prototype: a camera mounted on a rolling robot, running lightweight inference at the edge. Everything worked until I tried to line up video frames with sensor timestamps (IMU + wheel odometry) for clean motion tracking.

What broke was simple: my camera frames and my sensor streams weren’t on the same clock. On a dev laptop this kind of drift is annoying; on a robot it turns into jittery perception because frame-to-sensor alignment is off by tens of milliseconds.

So I went hunting for a practical way to synchronize clocks across the network using PTP (Precision Time Protocol) and then push that into a 5G/6G-integrated edge pipeline. PTP is a network protocol that distributes a high-precision time reference so multiple devices can share “time 0” to within microseconds (in the best case) instead of just “approximately”.

My specific niche goal: use PTP time as the timestamp source for frame alignment inside an edge video pipeline, while the network path runs over 5G/6G.


What I built (architecture in one picture)

I ended up with a pipeline like this:

  • PTP Grandmaster: a time server (or a master clock) reachable over the network.
  • Edge device (running the robot vision service):
    • joins the PTP domain (so its system clock matches the grandmaster)
    • captures camera frames
    • asks the OS clock for the timestamp at the moment the frame is taken
    • packages that timestamp alongside the frame for inference
  • Mobile/backhaul network (5G/6G):
    • only responsible for connectivity
    • the key is that the edge device’s clock is synchronized via PTP over the network

The important bit: I wasn’t trying to “timestamp the stream at the sender and hope.” I anchored timestamps to the edge device’s synchronized local clock, which PTP corrects.


First test: proving PTP was actually disciplining the clock

Before touching video, I verified time sync. On Linux-based edge devices, PTP is often implemented via ptp4l (from linuxptp). Here’s the check I used:

1) Install and run PTP (Linux)

sudo apt-get update sudo apt-get install -y linuxptp

Then, in one terminal, start the PTP client on the edge device. Replace eth0 with the actual interface connected to the network:

sudo ptp4l -f /etc/linuxptp/ptp4l.conf -i eth0

If you don’t have a config file yet, a minimal approach is to use command-line options, but I prefer config so I can pin the PTP domain and message interval consistently.

2) Confirm synchronization state

Another terminal:

sudo pmc -u -b 0 "GET GRANDMASTER_DATA_SET" sudo pmc -u -b 0 "GET TIME_STATUS_NP"

I looked specifically for “locked” / stable time status (the exact phrasing varies by build). The key outcome: the edge device’s time was no longer drifting freely.

Why this mattered: if PTP wasn’t locked, the whole timestamp pipeline would still be “almost correct,” which is worse than plainly incorrect because it can be subtle.


The video timestamp pipeline (frame-to-sensor alignment)

Next I coded a tiny edge service that:

  1. captures video frames
  2. attaches a PTP-synchronized timestamp
  3. emits a JSON line per frame (so I could compare to sensor timestamps)
  4. runs fast enough to keep up with a test camera

Step-by-step: a working Python capture service

This example uses OpenCV to grab frames and a monotonic clock that we tie to system time. PTP synchronizes system time, so we read the current time from the OS at frame capture.

Technical note:

  • system time = wall-clock time (disciplined by PTP)
  • monotonic time = never goes backwards (good for measuring durations)
    For alignment, I used system time because that’s what PTP updates.
import time import json import cv2 from datetime import datetime, timezone def utc_iso8601(ts_seconds: float) -> str: # Convert epoch seconds to an ISO-8601 UTC string for readable logs return datetime.fromtimestamp(ts_seconds, tz=timezone.utc).isoformat() def capture_frames(camera_index: int = 0, width: int = 640, height: int = 480): cap = cv2.VideoCapture(camera_index) cap.set(cv2.CAP_PROP_FRAME_WIDTH, width) cap.set(cv2.CAP_PROP_FRAME_HEIGHT, height) if not cap.isOpened(): raise RuntimeError("Could not open camera") frame_id = 0 while True: ok, frame = cap.read() if not ok: break # Timestamp at the moment after we successfully read the frame # (with PTP-disciplined system clock on this device) now = time.time() # seconds since epoch (system time) payload = { "frame_id": frame_id, "timestamp_utc": utc_iso8601(now), "timestamp_epoch_s": now, "shape": [int(frame.shape[0]), int(frame.shape[1]), int(frame.shape[2])] } # Output as JSON lines so the edge-to-network log pipeline can ingest it print(json.dumps(payload), flush=True) frame_id += 1 if __name__ == "__main__": capture_frames(camera_index=0)

What happens when I run it

python3 frame_capture_ptp.py

I saw lines like:

{"frame_id":0,"timestamp_utc":"2026-10-05T12:34:56.123456+00:00","timestamp_epoch_s":1765053296.123456,"shape":[480,640,3]} {"frame_id":1,"timestamp_utc":"2026-10-05T12:34:56.157890+00:00","timestamp_epoch_s":1765053296.15789,"shape":[480,640,3]}

Then I paired that output with IMU/wheel odometry logs that were timestamped using the same edge system clock.


The network part: tying it to 5G/6G edge deployment

The “5G/6G integration” piece is less about coding a radio driver and more about deployment reality:

  • the edge device moves between networks or has varying latency
  • you don’t want your perception pipeline’s correctness to depend on network jitter
  • you want time sync to survive the network environment

Here’s how I tested the connection impact without getting lost in radio complexity:

1) Put PTP client on the edge device over the 5G interface

On the edge device, PTP ran over the interface that carried traffic through the 5G modem.

2) Stream capture logs through the network

I used a simple TCP forwarder for logs. One terminal on the “receiver” side:

nc -l 9000 > ptp_frames.log

And on the “edge” side:

python3 frame_capture_ptp.py | nc <receiver-ip> 9000

The important thing I measured: timestamps stayed consistent even while the network path was noisy.

If you only look at frame intervals at the receiver, network buffering can smear timing. But because each frame carried a PTP-disciplined timestamp generated on the edge device, alignment remained meaningful.


The real gotcha I hit (and fixed)

The subtle bug: I initially compared IMU timestamps against camera timestamps using different clock sources.

  • IMU timestamps came from a monotonic clock (good for durations)
  • camera timestamps came from system time (PTP-disciplined)
  • my alignment code assumed both were epoch seconds

That produced “wrong by a constant” results and made it look like PTP was broken.

The fix: normalize everything to the same time base

I changed the sensor logger to also emit epoch seconds from system time, the same time.time() basis as the camera capture service.

Here’s the sensor-side pattern (minimal):

import time import json while True: imu_now = time.time() # system time aligned by PTP imu_payload = { "timestamp_utc": time.strftime("%Y-%m-%dT%H:%M:%S", time.gmtime(imu_now)), "timestamp_epoch_s": imu_now, "ax": 0.01, "ay": -0.02, "az": 0.98, } print(json.dumps(imu_payload), flush=True) time.sleep(0.01)

After both streams used the same clock base, frame-to-sensor delta calculations snapped into alignment.


Minimal alignment check: compute frame-to-IMU delta

Once both logs were collected, I ran a small script to compute time differences. I stored camera and IMU lines in separate files, then matched nearest timestamps.

import json from datetime import datetime, timezone def load_lines(path): with open(path, "r") as f: return [json.loads(line) for line in f if line.strip()] def nearest_imu(frame_ts, imu_list): # Simple nearest-neighbor by absolute time delta best = None best_abs = None for imu in imu_list: d = abs(imu["timestamp_epoch_s"] - frame_ts) if best_abs is None or d < best_abs: best_abs = d best = imu return best, best_abs def main(): frames = load_lines("ptp_frames.log") imus = load_lines("imu.log") imus = sorted(imus, key=lambda x: x["timestamp_epoch_s"]) # Pair first N frames for a quick sanity check for f in frames[:50]: imu, delta = nearest_imu(f["timestamp_epoch_s"], imus) print( f"frame_id={f['frame_id']} " f"frame_ts={f['timestamp_epoch_s']:.6f} " f"imu_ts={imu['timestamp_epoch_s']:.6f} " f"abs_delta_s={delta:.6f}" ) if __name__ == "__main__": main()

When PTP was locked and both streams used system time, the abs_delta_s numbers clustered tightly (in my test environment, on the order of a few milliseconds; the exact value depends heavily on camera buffering and driver timestamp behavior).


What I learned building this

I learned that “5G/6G integration” in a Physical AI robot isn’t primarily about clever networking code—it’s about making timing and inference pipelines resilient to network reality. By anchoring frame timestamps to a PTP-disciplined system clock on the edge device, the robot’s perception pipeline stayed consistent even when I stressed connectivity and log transport. Most importantly, I discovered that clock base mismatches (monotonic vs epoch/system) can masquerade as network or PTP problems—so I unified all timestamps to a single time base end-to-end.