Euan Cowie
§

Work

Shipping software to the North Sea

Platform engineering over a hybrid AWS and on-premise estate — Terraform, Kubernetes, iPXE, CI/CD — behind a GPU inference product on oil rigs and wind farms. Provisioning went from a week to under four hours.

OrganisationZelim
Period2023 — 2025
StackAWS / Terraform / iPXE / Edge / WebRTC / Hybrid cloud

The problem

For two years I was the platform engineer at Zelim: Infrastructure-as-Code over a hybrid AWS and on-premise estate, Kubernetes, CI/CD, the network, the release process, and the observability over all of it. The product on top was a GPU inference system doing real-time detection on video. Most of that is ordinary platform work, and the reason it is worth a page is the environment, which punished every assumption platform work usually rests on.

ZOE is a maritime detection product. It does not run in a datacentre. It runs on oil rigs, on offshore wind farms, on cruise ships and in harbours — installations that move, that lose connectivity for long stretches, and that you cannot reach except by helicopter or boat. It is also life-safety equipment, which means a regression that reaches the field is not a support ticket.

An interactive camera view from an offshore platform at night, drawn as glowing lines on a dark display. A small boat passes on a moonlit sea and the tracker picks it up. At twenty seconds someone goes over its side: the display raises a man overboard alert, flashes red, and turns the camera to them, then keeps them in view as they drift while the boat carries on.

An illustration, not ZOE's interface: a platform camera at night, a boat passing 110 m out, and at 20 s someone going over the side. Drag to look around, scroll or pinch to zoom, click a target to follow it.

That inverts most of what normal software delivery assumes. There is no "just redeploy". There is no shelling in to see what happened. A bad release is not an incident you roll back in four minutes; it is a vessel that has to finish a voyage before anyone can touch it.

When I joined, provisioning one unit took a week, deployment configuration and product code were tangled together so every install was bespoke, and the installs themselves were carried out by third-party contractors rather than by us.

Provisioning: a week to under four hours

The headline number is that bare-metal provisioning went from a week to under four hours. The mechanism is the more interesting half.

I built the iPXE boot process from the ground up: network-boot a target host, bake the OS onto it, and register it with AWS SSM. From there an SSM Document carries out the full product installation onto the selected host. No engineer walks a machine through setup.

But automating a bad process just produces bad results faster. The load-bearing change was separating deployment configuration out of the product, into its own repository, consuming a versioned schema published by the product monorepo and validating against it. A deployment that would not work is now rejected before it is attempted, rather than discovered on a vessel. Health checks and rollback sit on the deployment path behind that.

That is a release-engineering position rather than an automation one: treat deployment configuration as a versioned contract the product publishes, and validate at the boundary. Make the invalid state unrepresentable instead of catching it later.

The estate: genuinely hybrid, on purpose

Zelim ran across AWS and an on-premise colocation footprint at Pulsant South Gyle in Edinburgh, and both were permanently in use. I architected and led Infrastructure-as-Code across the company to cover both — Terraform with HCP for state and workspace management, alongside CDK and CDKTF, with siloed environments spanning the cloud and the data centre.

The result engineers cared about was self-service: software and ML engineers could stand up what they needed without queueing behind one person. The result the company cared about was a single central view of the whole footprint, across both substrates.

Pulsant published a case study on that decision in April 2025 — Revolutionising maritime safety with smarter systems — in which I am quoted on the requirements and the reasoning alongside Zelim's CTO. It is the closest thing to a third-party account of this work, and I would rather link it than paraphrase it.

It needs one correction, and I would give the same one in a room. It frames this as a move off the cloud. It was not. AWS kept carrying global deployment infrastructure and networking throughout — the page compresses a placement decision into a repatriation story, which is the story a colocation provider is in business to tell. The colocation footprint was stood up for two workloads it genuinely suited better:

Developer environments that were ground truth for production hardware. Engineers developed against the same physical hardware the software would be deployed on, rather than a cloud approximation of it. For GPU inference on specific edge devices, that is the difference between a test that means something and one that does not. I led the project to standardise those as hosted GPU VMs: 20+ developers on them, build times from over 40 minutes to under 10 via a shared host cache and the faster network inside the data centre, and the per-developer configuration burden gone entirely.

24/7 video inference for continuously-monitored sites such as wind farms, where the workload is constant rather than bursty and the cloud economics are the wrong shape.

So: a deliberate placement decision per workload, not a repatriation. The reasoning is the interesting part, and it is also why the IaC platform had to span both — because both were permanent.

Getting the packets there

Delivering software to the assets and watching them both assume you can reach them. I built the network backbone that made that true: secure remote access to every deployed site for observability, updates and video retrieval, with resilient failover using FortiGate HA paired to Transit Gateway VPN attachments and CloudWatch alerting, over cellular and satellite hardware.

15+ sites on it. I also led the research that proved remote ingestion and inference viable over cellular and Starlink links for wind farms, rigs and harbours — feasibility work that unlocked a product direction.

Observability ran on Prometheus and Grafana for the fleet with AWS SSM Agent and CloudWatch on the control-plane side, which is the split you want on a hybrid estate.

The video pipeline

The real-time stack was WebRTC and SRTP through an SFU, with NVIDIA hardware encode and decode and ffmpeg. The change that mattered was architectural rather than tuning: I re-architected the pipeline to restream, so a single ingest fans out instead of every consumer pulling its own stream. On bandwidth-starved satellite and cellular links, cutting network load and RTP packet loss that way is what makes everything upstream viable.

Alongside it, multi-architecture build pipelines with Docker buildx using NVIDIA Jetson hardware as native ARM capacity, which is what opened the Small Form Factor product line — a new commercial revenue line that existed because the builds did.

Underneath the maritime framing this is edge inference, and the problems are the generic ones: getting a model onto constrained ARM hardware, building for an architecture your CI does not natively run, keeping the container runtime and driver stack in agreement on a device nobody can log into, and deciding what inference happens on the device versus what is worth the bandwidth to send back. A phone, a kiosk or a factory camera poses the same questions with a shorter boat ride.

Delivery, and the humans in it

I directed the internal DevOps strategy: deployment frequency up 300%, onboarding from days to minutes. Self-hosted CI runners with a local Git LFS cache shared across the runner host, so repeated fetches of large binary objects were served locally rather than billed by the gigabyte. AWS CodeArtifact via OIDC federation, so CI authenticated with short-lived credentials rather than long-lived secrets.

Because the product is life-safety-critical, I introduced test parallelisation across ML, frontend and backend, and nightly runs performing safety checks on long-running soak tests. Most teams parallelise tests to save CI minutes. This was to stop a regression in something that can kill someone from reaching a deployment.

And the parts that are not software at all: I was technical lead for the ISO drop-test certification, which passed first time — something few achieve. I specified the complete deployment across both sides of the hardware line — server, disk array, iSCSI, Docker, NVIDIA container runtime, FortiGate firewall and switch, and the cabling — and wrote the installation guide and checklists third-party installers followed onto cruise ships and oil rigs. I led the technical integration meetings with customers ahead of those installs.

What it meant

Two things I took out of this.

The first is that when the cost of being wrong is a boat trip, you push correctness earlier — into images, into pipelines, into a schema at the deployment boundary — rather than relying on being able to intervene later. That habit transfers to anything with a long feedback loop, which is most things that matter.

The second is that almost everything here was built to be operated by someone else. The IaC platform existed so engineers could self-serve. The installation guides existed so contractors could install without me on the call. The DevOps and cloud architecture work was handed off to Ops to run. I also mentored the data team on building scalable data architecture in AWS. A system that only works while its author is watching it is not finished, and offshore is simply the environment that makes that obvious.


Elsewhere: Pulsant, Revolutionising maritime safety with smarter systems, 28 April 2025 — a client success story on the colocation decision, quoting me as Software Engineer at Zelim alongside CTO Doug Lothian. Read it with the correction above: it is a vendor's page, accurate about the relationship and shaped to flatter the vendor on the architecture. Some of what I describe in it was intended rather than done at the time of writing.