Euan Cowie
§

Work

A whole product, alone

A React Native app, a Neo4j food graph doing graph RAG with fail-closed safety, and a Kubernetes cluster under Argo CD — built and run by one person. Plus an offline sync engine, designed and measured against six alternatives before a line of it was written.

OrganisationHeilsa
Period2025 — present
StackReact Native / Neo4j / Graph RAG / Postgres / Kubernetes / Agents / Architecture

The problem

A React Native app on Expo, a Convex backend, a Spring Boot service over Neo4j doing graph RAG, and a Kubernetes cluster under Argo CD with the full observability stack in it — written and operated by one person. The domain is household food, which is the part that sounds trivial and is the reason the engineering is interesting.

Household food decisions are split across three places that never talk to each other: what is in the cupboard, what you were going to cook, and what you need to buy. Plenty of apps have tried to join them up, and they died the same death every time — keeping an inventory meant typing one in, so households quit within a fortnight.

The input cost is what killed the category. Multimodal models made passive capture viable at roughly the point I started, which is the only reason this was worth attempting again.

What it is

Heilsa is a shared pantry and meal planner for households, on iOS. I founded HEILSA LTD and built all of it myself: the app, the backend, the knowledge graph, the infrastructure and the marketing site. I am the only engineer and I write all of the code.

It is live on TestFlight with a small closed beta. My own five-person, multi-diet household has run entirely on it — pantry, meal planning, shopping — daily since February 2026, with eight external beta households onboarded. The traction claim I care about is retention depth rather than downloads: a product that only works if every member of a household keeps using it, still in daily use after six months.

What follows is the target architecture — where the system is going, laid out as a design exercise — with a clear account at the end of which parts are running today and which are next.

What I learned the hard way

The first version was offline-first on top of a backend that does not do offline, with a hand-rolled outbox replayed by a background worker. It ran in production and failed in four ways that were each fixable and together a signal: writes that failed terminally with no message, records stuck in states no code path could leave, local "transactions" that committed before their async body ran, and cached derived state that disagreed with its sources. None of it failed loudly, and the test suite stayed green throughout. I deleted it and wrote down the conditions under which offline would be allowed back.

Those conditions are the spine of the design below: real transactions at both ends, a mutation contract where an omitted field is a build error rather than a silent drop, and a route for a permanently-failed write to reach the person who made it. A sync engine that cannot satisfy all three is a queue with better marketing.

The shape of the system

Three planes and a set of managed edges. The device holds a real database and a rebase engine, so the app works with no signal. The data plane is Supabase in the EU: Postgres is the durable authority for everything a household says about itself. The compute plane is a Kubernetes cluster I run on a netcup VPS in Germany: the graph, the agents, the sync endpoint and the platform underneath them. The edges — Resend, Bedrock, Transcribe, S3 — are the managed pieces not worth self-hosting on a cheap cluster, and each is reached only from the compute plane.

The boundary rule that makes the rest coherent: each fact has exactly one authoritative owner, and everything else is a projection. Postgres owns what a household has, wants and cannot eat. The graph owns what is true about food. Agents own nothing — they propose.

What it promises a household

Before the parts, the reasons. Everything below exists to deliver six guarantees to the people using it, and each one is something the apps in this space either cannot offer or do not:

  1. It works in the aisle. Reads and writes with no signal — the pantry, the list, the plan and your recipes — in real transactions, and a write the server will not accept comes back to you as an undo, never as silence.
  2. Nothing is lost and anything can be undone. The pantry is a ledger, so undo is a mechanism rather than a per-feature afterthought, and waste is distinguishable from eating.
  3. It is safe for the person with the allergy. Recommendations fail closed: a recipe the graph cannot prove safe is not shown, and every exclusion can be explained.
  4. It never writes into your kitchen unasked. Capture and agents produce candidates for review. Nothing a model says reaches shared household state without a person accepting it.
  5. It knows your kitchen, not a generic one. The agents are grounded in your household's own consumption history and the graph, so "what can I cook tonight" is answered from what is actually on the shelf.
  6. Your data stays in Europe, on infrastructure one person controls. Swiss and German hosting, EU model endpoints that do not train on inputs, and no consumer-data pipeline to a US cloud.

Offline you can trust

The engine is the design Replicache, Zero and Linear converged on independently — a server-authoritative mutation log with rebase — implemented at this system's scale behind a three-function port, push, pull and poke, so the backend is an adapter and nothing above the port knows which one exists. It is built rather than bought, and this time the reason is a measurement rather than a preference: before committing, I built a simulator, ran six candidate designs through it, and chose from the numbers.

Six designs, one harness

The simulator is a seeded household of three members, a network that delays, drops and duplicates, one server with a handler per protocol, and six client strategies sharing the same domain and the same planner: today's online-only shape; this engine with whole-scope snapshot pulls; the same with row-version diff pulls; PowerSync as documented, with plain rows; PowerSync with append-only shopping events; and a hybrid with this engine's command log on the write path and PowerSync's buckets on the read path. Eight scenarios — a supermarket basement, cooking on lossy Wi-Fi, a phone killed mid-upload, a member revoked while offline, conflicts plus a schema bump, a growth sweep to a ten-thousand-recipe catalogue and a two-thousand-recipe library, three days offline against one day of deletion retention, and a stress with a quarter of messages dropped and a third duplicated — each with poison commands and transient server errors, twenty seeds each. Every design converges with zero silent failures; what separates them is cost and blast radius.

  • Whole-scope snapshot pulls do not survive a growing library: 32 MB per device per day at a twenty-recipe library, 200–280 MB at two thousand. Diff pulls and buckets sit at 0.4–3 MB, within five percent of each other.
  • Plain rows lose shopping updates — four per run in the aisle scenario, fourteen at worst — because an in-place counter carries no intent. Events and commands lose none.
  • A write the server cannot process costs one told rejection under commands. Under a row upload it blocks the queue, holds downloads back for the retry window and takes the rest of its batch with it when discarded.
  • Diff pulls and the hybrid are indistinguishable on every safety and cost number. The only difference is whether the read path is owned or bought, so the engine owns it, and there is no sync service to run.

The simulator also found a bug in the design as I had written it. After a crash between a rejection and its acknowledgement, the phone replays the mutation, and a server that skips seq ≤ watermark with a blanket "ok" turns that rejection into a silent success. The server now stores every mutation's outcome and re-answers a replay with it. That is exactly the class of thing example-based tests never find, and the reason the simulator was the condition rather than a nice-to-have.

Why not buy

PowerSync is the mature thing to buy, and the numbers now say what the argument used to. It syncs Postgres to on-device SQLite by bucket and replays row diffs through an upload queue; Heilsa wants commands. The pantry is already event-sourced at the data level — a movement is an event, and undo is a contra movement — so a command log is not extra machinery here; it is the domain's native shape, and it is what makes rejection-with-undo and replay cheap. The read set stays bounded by design rather than by hope: a household syncs its pantry projection and a window of its ledger, its list and plan, and its library — the recipes it wrote and copies of the curated ones it saved — while the catalogue itself is browsed online and copied on save, so a ten-thousand-recipe catalogue costs a phone exactly what a hundred-recipe one does.

Zero was online-first with shallow offline writes when I surveyed it. LiveStore is the closest philosophical match — an event log with materialisers — and was a beta, and betting a one-person product on a beta sync engine is how you get a second removal document.

The three conditions, made concrete

Real transactions at both ends. On the device, appending a command to the pending log is one SQLite transaction, and the engine takes now() and random() as injected dependencies so a simulator can drive it. On the server, the push endpoint runs in the household-workflow service on the cluster, and per mutation it does exactly this: BEGIN → lock the household row with SELECT … FOR UPDATE, which serialises that household's writes without global serializable isolation → an outcome already stored for this mutation? re-answer it → seq ≤ watermark and unseen? reject it as out of order → run the planner against canonical state → apply, the movement and the projected quantity together → scope_versions++ → store the outcome → watermark ← seq → COMMIT → acknowledge. Three small tables carry exactly-once over an at-least-once network: client_watermarks, mutation_outcomes and scope_versions. This is the concrete reason for Postgres — I need a transaction I can trust with several writes.

That service is TypeScript, and I want that decision visible rather than buried: the same planner has to run on the phone and on the server, and porting it to Go would break that guarantee silently. It is a deliberate exception to the Go-and-Bazel direction for services, and it is written down as one.

A contract where an omitted field is a build error. Commands are typed domain intents — pantry.addItem, not a row diff — validated by the same schema at both boundaries by the same code. There is no serialiser to forget a field in, because there is no invented wire format. The schema version travels with every push and pull, so a stale bundle gets a gate rather than a silent drop.

A route for a failed write to reach the person. Every mutation is acknowledged individually and its outcome is kept. A rejection advances the watermark past it, the client drops it, the next rebase makes the speculative effect vanish, and the UI shows an undo toast built from the command payload itself. Rejections never block the queue. The failure that was invisible last time is now a notification with a button.

Pull, poke and auth

Pull goes straight to Postgres: a SECURITY DEFINER function called through PostgREST. Not a member → forbidden. Version unchanged → notModified. A first sync, or a phone further behind than the retained change index → a snapshot of the scope. Otherwise a diff: every row of the scope with a version newer than the phone's, and the ids deleted since. Every synced table carries a row version, and a change index and a deletion log keep seven days each, so the phone never has to remember what vanished. Row-level security sits underneath as the safety net, but scope evaluation is what actually implements revocation — a row policy cannot reach into a phone's disk, and a pull that returns forbidden can, because the device deletes the scope. The poke is Realtime change-data-capture on scope_versions: a hint, never a data channel. Auth is a Supabase JWT the workflow service verifies against Supabase's JWKS. Storage holds photos; edge functions receive webhooks and nothing else.

The ledger over time

The pantry ledger is append-only, and a phone cannot hold it forever, so it is three layers. The projection — the current quantity on each lot — is written in the same transaction as the movement, and the planner reads it and never folds. The window — movements newer than ninety days — is what the phone holds for undo and the activity feed, dropped by a nightly job that uses the server's clock rather than the device's. The archive is everything older, on the server only, fetched on demand for history. Window membership is a function of time, so expiry needs no deletions; undo is offered for what the phone holds, and the server, which holds everything, accepts an undo at the window's edge that the phone can still see. The overnight reconciliation keeps asserting that projection equals the fold of the whole ledger, which is what makes the projection safe to plan against.

That design was simulated too, over a year, with one phone off for a hundred days and another for fourteen and with undo aimed at and past the window's edge: every device converged, nothing was silent, the projection never drifted from the ledger, the nightly job never moved a balance, the phone held about 2,600 movements against 10,000 on the server, and every undo of an archived movement was refused on the spot. The projection costs about one percent on the wire.

What it deliberately is not

No CRDTs — nothing here is multi-writer on the same field, and quantities are movements, not counters. No tombstones on the phone — the server keeps a deletion log with seven days of retention, and a phone further behind than that gets a snapshot. No wire format. No merge tables — the planner's ordinary semantics decide conflicts, and if that ever fails the escalation is per-property last-writer-wins, not a CRDT. No sync service.

Three things would make me stop and buy: a household's synced set passing ten thousand rows or ten megabytes on first sync, a simulator scenario the engine cannot be made to pass, or the engine becoming a bus-factor risk because nobody but me can touch it. The port is exactly where PowerSync would drop in — the simulator's hybrid strategy is that swap, and it shares the write path line for line — so the switch costs the read adapter and nothing above it.

The pantry is a ledger, not a list

This part is built and running, and it survives the move to Postgres unchanged because it was designed against a property of transactions, not of a vendor.

A pantry stores lots — one quantity of one thing, acquired once. Store only the current state and every write throws away what you would need to undo it, so every undo becomes its own feature, built afterwards out of whatever the write happened to leave behind. That bit me in production: a slot marked cooked by accident could have its plan restored but not its food, because the food's history was gone.

The model underneath is inventory accounting, and the fix is the one inventory accounting settled on two centuries ago: keep the movements, derive the balance.

The journal is authoritative and append-only; the current-quantity table is a projection of it, maintained in the same transaction that posts the movement, so the two cannot drift on partial failure. The only real drift risk is a write path that bypasses the journal, which is why there is exactly one seam that can post one. A lot is identified by the id of the movement that opened it — self-referential, stamped immediately after insert, because a document cannot know its own id before it exists.

Four things fall out of one mechanism rather than four features: undo for any pantry write, waste distinguishable from eating, real turnover rates read straight off the movements, and a household audit trail. The third of those is what the agents are grounded in.

What is true about food

Postgres owns what a household says. A separate Spring Boot service over Neo4j owns what is true about food: canonical ingredients, allergens, food taxonomies, substitutions, the curated recipe graph, and whether a given recipe is safe under a given set of constraints. It is queried from the cluster only. No device ever reaches it, and the graph is never replicated to a phone — it is large, shared across every household, and read-only from a client's point of view. What a household saves is different: a saved recipe is copied into that household's own library, which is a sync scope like the pantry, so the recipes you cook from are on the phone and the catalogue you browse is not.

Recommendations fail closed. A request carries the household's current constraints. The service scopes candidates to what that user may see — curated public recipes, their household's, their own private ones — then excludes anything inactive, incomplete, stale, unresolved or allergen-conflicting before it ranks what is left. It returns the recipes, the safety facts, the reasons, and two version stamps: the preference version the answer was computed against and the graph version it was computed on. The caller stores the result only if the preference version is still current, and hides anything whose safety facts are missing or stale. A recipe the graph cannot vouch for is not shown, and every exclusion has an answer to why.

Graph RAG, in the literal sense. Retrieval is a two-stage thing: a vector index over ingredients and recipes finds candidates by meaning — the embedding index runs inside the same service — and then the graph does what vectors cannot, walking substitution, allergen and food-group edges so that "what can I cook tonight" is answered against the household's actual shelf and the constraints of everyone eating. The model provider proposes ingredient matches with provenance and a confidence; it is never the final authority for safety, and changes to the canonical graph go through a reviewer.

I am not going to describe how ingredient identity actually resolves — turning "tinned chopped tomatoes", a recipe calling for passata and a supermarket listing into the same underlying thing across diets and substitutions is the part that took longest and is the part worth protecting. The principle is publishable because it is a stance, not a trick: it is deterministic, and an incorrect canonical ingredient is worse than an unresolved one. It would rather hand you a suggestion to confirm than quietly link the wrong food into a kitchen where someone has an allergy. Canonical identity is derived on the trusted save path, never by the parser that produced the text.

Agents that ask first

Three agents run on the cluster. They share a rule that is the whole design: an agent may read the ledger and the graph, and may write only candidates. Household state changes when a person accepts a candidate, through the same single write seam as every other mutation. A model's output is a proposal with provenance, never a fact.

Capture turns a photo, a receipt, a typed sentence or a spoken one into structured pantry rows and reads a recipe's ingredients, steps, timing and servings off a few photos of a page. Voice arrives through streaming transcription. Every row lands as a candidate; accepted, edited and rejected rows are tracked, which makes the evaluation signal for extraction quality a by-product of normal use rather than a separate harness.

Planning drafts the week. It reads turnover from the ledger, the plan slots already chosen, and the household's constraints, then asks the graph for safe candidates and proposes a plan. It never writes a slot.

Steward is the quiet one: what is running low given how fast this household actually uses it, what is about to expire and what could be cooked from it, what the shopping list is missing for the plan. It produces suggestions and the occasional email digest, and it is the agent that turns the ledger's turnover data into something a household can feel.

All three call the model provider through Bedrock in an EU region, chosen for two properties: model choice without vendor lock, and a contractual guarantee that inputs are not used for training. Receipts and household photos are exactly the data that should not leak into someone's next model.

The cluster, on a cheap box

The compute plane is a Kubernetes cluster on a netcup VPS in Germany — one node today, deliberately, with the storage class and the runbooks written for the day it becomes three. Argo CD owns everything inside the frame from a Git repository that is separate from product source: bootstrap, then infrastructure, then apps, each a root application. Cilium for networking and policy, ingress-nginx and cert-manager for the edge, and a full observability stack in-cluster, with Sentry and PostHog on the app side.

Every service is built from a Bazel monorepo with versioned protobuf contracts and follows one runtime contract: liveness, readiness and version endpoints, graceful shutdown, non-root and read-only filesystem, dropped capabilities, resource limits, and digest-pinned images so that what runs is exactly what was reviewed. Services are decomposed by fact ownership, not by screen or table — the household workflow stays one modular service rather than being split into pantry, shopping and recipes prematurely, and specialists are extracted when they have their own source of truth, their own scaling shape, or retryable async work: the graph, the agents, notifications.

One exception is drawn on the diagram because it is the most important line on it. Host-level survival tooling sits outside Argo CD. GitOps that depends on a healthy Kubernetes API is exactly what you cannot use when the Kubernetes API is the thing that is broken. The etcd backup, and the runbook to restore from it, live on the host and ship off-site to S3.

What makes this cheap is not the provider. It is that nothing here needs a managed control plane, a load balancer per service, or a GPU: inference is bought per token from an EU endpoint, storage that must survive the box is off-site, and everything stateful that is not the graph lives in Supabase.

The managed edges

Each of these is used for one thing, chosen because self-hosting it on a small cluster is the wrong trade.

Supabase — Postgres as the durable authority, row-level security as the first line of household isolation, Auth for accounts, Realtime for the poke, Storage for photos, point-in-time recovery with a nightly export to S3. Edge functions receive webhooks and nothing else.

Resend — every email the product sends: one-time sign-in codes with an escalating resend cooldown, household invites, the steward's digests. Templates live in the same monorepo as the app, and delivery events come back over a webhook to the notification dispatcher, which owns the fact of whether a message was sent — not whether the thing it was about happened.

AWS, narrowly. Bedrock in an EU region for every model call, for the two properties above. Transcribe streaming for spoken capture. S3 in the EU as the off-site target for every backup — etcd, the graph, the Postgres export — because a backup on the same provider as the thing it protects is a copy, not a backup. That is the whole footprint. No SES, because Resend. No managed Kubernetes, because the cluster is the cheap part. No RDS, because the database is the one thing worth paying someone else to run well.

Today and next

Honesty about the line between built and designed, because the two read differently to an engineer:

Running. The app on TestFlight, in daily household use. The pantry ledger. The Food Graph service with its embedding index, resolver and fail-closed recommendations. Review-before-save capture. The cluster under Argo CD with the full platform and observability layers, the service template and the etcd runbook. Resend for sign-in. The sync simulator, with six designs compared and the ledger design proven over a simulated year.

Next, in order. The sync engine against the simulator, with Postgres as the durable authority and its Supabase adapter — the design above, specified down to the table schemas. Then the household-workflow service on the cluster, in TypeScript for the planners' sake. Then the agents, then Transcribe for voice. Bedrock replaces the current model provider when the agents land.

What it meant

I have kept it deliberately small — self-funded, no employees, no outside money, not hiring, not growing. It is a large hobby rather than a business straining to escape, and it is how I keep my hands on the whole stack: the mobile UI, the sync engine, the data model, the graph, the cluster, and the dashboards that tell me when any of it breaks.

The thing I would want an engineer to take from this page is not any single component. It is that every part is justified by a promise to the household — offline that surfaces its failures, a pantry that can be undone, safety that fails closed, models that ask first, data that stays close — and that the system got there by building the ambitious version, measuring how it failed, and designing the second version against the specific ways the first one lied.

It is installable. If that is the question you are asking, take it for a run rather than taking my word for it — heilsa.io has a join page for the beta, and I hand out TestFlight invites myself.


Elsewhere: heilsa.io — the product site, with a public join page for the closed beta.