Long-Distance Relationship: You and Your LLM — the Slides, and What's Coming
Article

Long-Distance Relationship: You and Your LLM — the Slides, and What's Coming

I gave this talk today at Future Coding Day 2026 in Copenhagen. Here's the one-sentence version, the slide deck to download, two other talks from the day worth your time, and where the fuller write-ups are headed next.

I gave a talk today at Future Coding Day 2026 in Copenhagen. Thirty minutes, including questions — which isn’t enough time to argue a position properly. It’s enough time to state it once and back it with a number nobody in the room can wave away.

The title was Long-Distance Relationship: You and Your LLM, and it starts from three real problems with running your organisation’s critical work on a frontier model sitting in someone else’s datacentre — and not just someone else’s: on the other side of the globe, in the US or China, under a legal system you don’t vote in.

Money. Two Claude Code sessions burned through €170 of expiring credits in three hours on a Friday — nothing unusual was happening, the meter just wasn’t being watched. Data. An agent that decompiles a licensed package to work out an undocumented API sends that decompiled source out with the prompt, and “we don’t train on your data” is a different promise from “we don’t keep it” — a court can override a vendor’s own retention policy regardless of what the vendor intended, and one already has. Supply. In June this year, a US export-control order suspended two frontier models for every customer on the planet, without warning, for eighteen days. None of that is an argument that any vendor did something wrong. It’s what happens when the infrastructure your work depends on sits abroad, and you’ve built your workflow around it as if that were permanent.

None of which means the answer is “stop using frontier models.” It means knowing what each task actually needs, instead of reaching for the most expensive option out of habit.

Capacity is not capability

The one-sentence version is on the slide behind me in the photo above — the deck’s own words for it, not mine:

Model selection is an architecture decision, not a preference.

The unit you’re actually selecting is model plus harness plus context, not model. You’re not choosing a brain, you’re choosing a body for it. What makes a model useful for a given task isn’t its parameter count — it’s how well it was trained, what harness it’s running inside, and what skills, tools and MCPs it has access to. Some open-weight models have gotten genuinely good, and I measured that rather than asserted it: the same pinned-spec task, a Python ray tracer with five milestones, run through Opus in a datacentre and qwen3.8:27b on my own GPUs, same harness throughout.

Opus (frontier, datacentre)qwen3.8:27b (local, vLLM)qwen3.8:27b (local, Ollama)
Time to finish27 min1h 502h 38
Render time at milestone 518.7s23.6s48.7s
Source lines written699420467
OutputIdenticalIdenticalIdentical

Not just the same picture — three different arithmetic routes converged on the same SHA-256 hash. Slower, and the local runs wrote shorter code doing it — but on consumer GPUs I own outright, running unattended overnight, with nobody watching and nobody needing to. Neither local run once emailed me about a reached token limit, the way a hosted plan will interrupt you mid-task. Slower doesn’t matter when nobody’s waiting — it’s the €170 story from the opening, seen from the other side.

It also matters because most of an agentic session isn’t the hard part to begin with. I once pulled apart a real troubleshooting session, start to finish: forty tool calls over twelve minutes to fix one production bug. Thirty of them were glob, grep and read — just looking for the problem. Nine ran the build to check a fix. One was the actual edit. Out of forty tool calls, one needed anything you’d call intelligence; the rest was mechanical legwork a well-specified model can do just as well, frontier or not.

What actually goes where

The practical split I’d argue for: run open-weight models locally, overnight, for whatever nobody needs to sit and watch — autonomous agents or batch jobs. Documenting dependency code. Building a RAG index of your own codebase. Running release-gate checks before anything is allowed to merge. Keep the frontier model for planning and specifying the work in the first place, which is the part that genuinely benefits from more capability rather than just more parameters.

Grab the slides

The full deck — 38 slides, the benchmark table, the “day shift, night shift” argument, and the playbook at the end — is downloadable as a PDF: Long-Distance Relationship: You and Your LLM (PDF).

Two talks that lined up with mine

Future Coding Day is the newer, more focused sibling to Future Product Days, and the two together were worth every minute — strong talks throughout, good conversations between sessions, and more genuine inspiration than I walked in expecting. Coding Day itself runs a tight programme — engineers, CTOs and founders building agentic systems, no sponsored talks — and two sessions on the day especially landed close enough to mine that I’d send you to them directly.

Matthias Lau’s The Hidden Economics of Agentic Software Engineering (Heureka Labs) followed mine on the schedule, and it was the right talk to follow it with. Mine argues that model choice is architecture. His goes straight into the economics of that architecture: dynamic model routing, why a static “always reach for the frontier model” choice stops scaling once you’re running parallel agent workflows, and the open design questions around when a routing decision should actually adapt. If you only watch one other talk from the day, watch that one.

Paul Stack’s When AI Writes the Code (Swamp Club) makes a related but distinct case: replace pull requests with issues, let the design constraints do the work up front, and treat “sharp constraints produce good code” as the actual mechanism rather than a slogan. What stuck with me was the gating — checks and adversarial-agent review running before anything is allowed to merge. That’s precisely the release-gate work I’d argue belongs on a local model rather than a frontier one. His talk and mine arrived at the same shape of answer from two different directions, on the same day, without either of us knowing what the other was building.

What’s coming next

Thirty minutes covers the shape of an argument, not the evidence for it. Over the next few weeks I’m writing up the parts that didn’t fit on stage: the full benchmark and its methodology, the case for right-sizing the model to the task, what actually leaves your machine in a prompt, why a cloud LLM is a jurisdiction risk and not just a vendor risk, what a night’s worth of GPU time is actually worth, and a couple of things that only made sense to write once I’d stopped worrying about looking dignified on stage. I’ll link each one here as it publishes.

If you were in the room today — thank you for coming, and for the questions afterwards. If you weren’t, the slides are above, and the rest is on its way.