Half a million genomes on a gaming GPU
A React and Rust/gRPC platform over an eight-GPU box, split from a monolith into services, with self-hosted observability — above a CUDA kernel that runs half-million-sample studies by keeping the cohort on the cards.
The problem
For two years I built and ran a distributed compute platform: React portals for researchers, a Rust and C++ control plane behind gRPC, a columnar data path in Apache Arrow and Parquet, and a self-hosted observability stack over all of it. Ordinary distributed-systems work. What made it unusual was the workload underneath and one constraint that ruled out most of the obvious answers.
A genome-wide association study looks for statistical links between genetic variants and a trait, across a very large number of variants and a very large number of people. On cohorts of 500,000+ samples, the conventional answer is cloud-scale compute: rent a lot of machines, submit the job, come back later. The shape of the science bends around that wait, and around that bill.
Omecu was commercialising work out of a University of Edinburgh spin-out that made those studies near-instant. Two questions followed from that, and I worked on both: could it run on hardware you can buy in a shop, and could the system around it be fast enough that a researcher explores rather than submits?
The kernel
I helped engineer and scale the "Golem Engine" — a custom GWAS kernel in CUDA. The constraint that shaped it was memory: a consumer NVIDIA 1080 Ti has 11 GB of VRAM, and no single card holds a 500,000-sample cohort. Copying the cohort in for every analysis would spend the run on the bus, so it is packed small enough to live across the eight cards in the machine, loaded once, and each analysis streams through it where it sits. The engineering is in the memory access patterns — keeping the device fed, keeping VRAM under the limit, and paying for the one transfer you cannot avoid once, not on every run.
Nearly half a million genomes on eight 11 GB cards
Ready · plays in real time
452,264 people × 623,944 variants
- Genotype
- 2 bits, packed
- Together
- 70.5 GB
- Kept
- 8.8 GB a card
- Genotypes
- 70.5 GB, packed
- On the cards
- 0 GB · 0%
- Loaded
- Once, then kept
PCIe 3.0: a ×16 link to each pair of cards, 12 GB/s, shared
- Each card holds
- 77,993 variants
- Load
- 1.5 s at 6 GB/s, once
- Each run
- 130 ms
- Tested
- 0 of 623,944
- Loading
- 0.0 of 1.5 s
- Top hit
- —
- Loci past 5 × 10⁻⁸
- —
- Inflation λ
- —
On eight GTX 1080 Tis, modelled, after loading the cohort once in 1.5 s. The cards’ arithmetic sets the pace. Omecu’s public figure for a whole GWAS was one or two seconds.
The result is that it processes 500,000+ sample cohorts on consumer hardware, outperforming cloud equivalents. That is the least replicable thing I have worked on, and to be precise about my share: the scientific core existed before me, and the kernel was a team effort I contributed to and helped scale rather than authored alone.
Data that cannot leave
Genomic data cannot be exposed to the network. That is a hard regulatory and ethical constraint, not a preference, and it shaped the architecture more than performance did.
I architected a two-container security model to guarantee data sovereignty, with a secure IPC protocol over Unix sockets streaming encrypted results between them. Raw data never reaches the network; results do, encrypted. The isolation is structural rather than policy-based, which is the only version that survives an audit question.
The same constraint ruled out every hosted monitoring platform, which turned observability into a build rather than a purchase.
The shape of the platform
Four planes and one rule. Researchers see statistics. A control plane dispatches work and remembers what it did. A results plane holds what came back. And the data holder — the only place a genotype exists — opens every connection outward and accepts none.
That last clause is most of the architecture. A site that dials out cannot be dialled into: there is no port to attack, no inbound exception to negotiate with a hospital's network team, and no configuration in which a query reaches the cohort by a route the site did not open itself. Inside the node, one agent holds a long-lived gRPC stream to the control plane and a Flight connection to the results plane, and it is the only process there with a network at all. The engine cannot send anything anywhere. It can only hand the agent finished statistics.
Inside the node, the engine and the service container are separate processes that share a Unix socket and nothing else. The engine holds the packed cohort across its cards and makes one pass per analysis. Phenotypes and covariates stay on the node as columnar files, read per job; the genotypes never move again after the first load.
The platform, with work crossing it
Measuring this browser
Five traces cross the same map.
Measuring what this browser does with a result, so the map can price its own wires.
- a wire — it shows its protocol and format while it is busy
- resident — loaded once at start-up, and never travels again
- a payload, drawn as thick as the bytes it is carrying
- accent — it is crossing the line around a data holder
- a box at work — the bar fills as its pass over the cards runs
- stopped — it asked for something no wire on the map does
Press a box to see inside it, and what its wires carry. Only three wires cross the line around a data holder; the three questions above ask for things no wire does at all.
The split is by what fails, not by what the code looked like. A gateway that speaks gRPC to services, gRPC-web to browsers and Flight to notebooks. An identity service, because a site must prove what it is before anyone pushes it work. A job service where every job is keyed by the hash of its query, so a retry cannot double-count and an identical question is answered from storage. A scheduler that knows what every device in the fleet measured when it joined, and a catalog holding metadata and nothing else — which cohorts exist, who may use them, what each job did. Splitting it that way is what let a GPU-bound path, a CPU-bound path and a serving path be scaled and reasoned about separately, instead of sharing a process because they started life together.
The data path is columnar the whole way. A result is numeric columns — an
effect size, a standard error and a p-value for each of a few million variants —
and once the arithmetic takes a fraction of a second, encoding those columns as
text costs more than computing them. So results travel as Apache Arrow
batches: the same layout in the engine, in the service container, on the wire,
and in the browser drawing the plot, with no parse step anywhere between. What
settles in the results plane is Parquet — row groups by chromosome, floats
in BYTE_STREAM_SPLIT and ZSTD, a page index and a Bloom filter on the variant
— which means a result can be read back a column at a time, months later,
without running anything again. Asking for one variant across every trait ever
computed becomes a file read, not a job.
Nothing assumes a card. The engine had to run on more than one kind of GPU, so the kernels came in variants chosen by what the device reported it could actually do, rather than one binary that assumed a 1080 Ti. The same principle runs up the stack: the scheduler sizes each device's share by the pass it measured when it joined, not by its name, so a node with newer cards takes more of the cohort without anyone editing a config file.
Two holders, one question. Cohorts that may not be pooled can still be asked the same thing. Each site runs the analysis over its own people and returns statistics only; the results plane combines them, inverse-variance weighted. Nothing is pooled, nothing is copied, and the answer is the one a pooled analysis would have given.
The cloud side stays deliberately thin. It dispatches, it remembers, it serves. It cannot reach into a node, and it holds nothing a data holder would mind losing.
Watching eight GPUs you cannot phone home about
The compute node ran 8× GTX 1080 Ti — 88 GB of aggregate VRAM, but only 11 GB available to any individual device. That asymmetry is the whole operational problem. Total GPU utilisation tells you nothing useful; what you need to know is whether one device is about to hit its ceiling, whether the workload is balanced across the eight, and which job caused a spike.
So I built a fully self-hosted observability stack:
- NVIDIA DCGM Exporter with Prometheus for per-GPU utilisation, VRAM consumption, temperature and workload behaviour across all eight devices.
- node_exporter and cAdvisor for host and container CPU, memory and disk.
- Grafana dashboards exposing VRAM high-water marks, GPU saturation, thermal behaviour and imbalance between devices during a run, rather than inferring them afterwards from profiling.
- Loki and Promtail for centralised logs, with Alertmanager handling thresholds — all inside the controlled environment.
Where the ecosystem did not cleanly give me application-level tracing across a Rust/C++/gRPC control plane, I hand-rolled it: propagating job and request identifiers across service boundaries and emitting structured logs, so a single GWAS job could be followed from dispatch through execution to result delivery.
The point was never "we used Prometheus". It was that a sovereignty requirement removed the option of buying the answer, and the workload had a failure mode — per-device memory pressure — that generic monitoring would not have surfaced.
This is the same problem as running training jobs on a shared GPU host, and it generalises directly. Aggregate utilisation is the metric everyone reaches for and the one that tells you least. What you need is per-device VRAM headroom, imbalance across devices, and attribution back to the job that caused a spike — whether the job is a GWAS, a fine-tune or a batch inference run.
The platform around it
I designed and built the rest end to end: the React client portals the researchers actually used, the control plane behind them, and the browser-based test automation with Playwright and Selenium that kept them honest.
Then I led the migration from a monolith to the service-oriented architecture above, which is what let the near-instant GWAS product launch commercially. It is one thing to draw that split and another to perform it on a system already in front of customers.
I also spearheaded the architectural decisions across the platform, and mentored two quite different groups: the junior engineers, on software engineering practice; and the academic and research staff, who wrote code but had never been trained as engineers. The second is the harder problem, and a different one — bringing researchers up to engineering practice is not the same job as mentoring someone who already knows they are an engineer.
What it meant
The lesson I keep from this: when compute gets fast, the bottleneck moves somewhere you were not looking. The interesting engineering was rarely in the maths. It was in the memory access patterns, in the boundaries between services, in the format results travel in, and in making a memory ceiling visible on a dashboard before it became a failed run.
It is also the clearest case I have of the same person writing the CUDA kernel and the researcher-facing UI on the same project — which turns out to matter, because the decisions at each end constrain the other.