I provision Linux servers. Not as a hobby exactly, and definitely not as a job title, as a consequence of wanting to know what I was designing on top of.
The machine is called olajidai. It runs Ubuntu Server on consumer hardware, built from bare metal: operating system and drivers, key-based SSH hardened onto a non-default port, static networking through nmcli and Netplan, remote access over Tailscale and SSH tunnelling rather than anything facing the open internet. On top of it sits a Docker Compose stack I assembled and maintain, n8n for orchestration, Hermes running as the agent on top of it, Ollama serving local quantized models, Open WebUI, Whisper for speech-to-text, Kokoro for text-to-speech, MinIO for storage, Caddy in front, Portainer for management, and a second GPU node running ComfyUI for image generation. The n8n image is a custom multi-stage build, because I needed a full FFmpeg inside a hardened container and nothing published had one.
It runs a content pipeline that takes a row in a spreadsheet and returns a finished, captioned, uploaded video with no human editing at any stage.
None of that is the interesting part. The interesting part is what it did to how I design AI features.
Latency is a budget, not a spinner
When you consume a hosted model, latency is a number you wait for and a loading state you design. When you serve the model, latency is a resource allocation you made. The GPU is either doing this or doing that. A larger model is a real decision with a real cost, paid in seconds, per call, forever.
That reframe has changed more of my design work than anything else on this list. I no longer design a loading state and consider the problem handled. I ask what we are spending the wait on, whether the user gets anything back during it, and whether the expensive call is even the right call, because I have sat and watched a queue back up behind a model that was one size too ambitious for the job it was doing.
A spinner is what you design when latency is somebody else's problem.
Failure is not an edge case, it is a schedule
In a demo, the model works. In production, at three in the morning, on a queue that has to be finished by breakfast, the model returns something structurally wrong, or the disk fills up, or an API token expires, or one service in the chain restarts and everything downstream of it silently receives nothing.
I have designed error states for years. I designed them differently after operating a pipeline that had to survive its own failures unattended. The questions became concrete: what does this step do when the previous one returns garbage rather than nothing? Is a partial result worse than no result? What is safe to retry, and what will quietly double-charge or double-post if you retry it?
Most AI product design treats failure as a single state, an error message, a try again button. Real pipelines fail in shades, and the design difference between a system people trust and one they babysit is almost entirely in how the shades are handled.
Local models make you honest about what the task needs
When every call goes to the largest available frontier model, you never find out how much of the intelligence you were actually using. Running quantized models locally forces the question constantly, because the capability ceiling is right there and you hit it.
And the answer, unglamorously, is that a lot of the work does not need much. Classification, extraction, reformatting, routing, tidying a transcript, small local models do these well enough that reaching for something bigger is a habit rather than a requirement. What genuinely needs the frontier model is a much shorter list than I would have guessed before I had to allocate the compute myself.
That maps directly onto product design. Feature ideas that begin with what if the AI could are usually asking for judgement. Feature ideas that begin with the user has to do this tedious thing every time are usually asking for structure. The second category is where most of the durable value is, and it is systematically underrated because it does not demo as well.
Orchestration is the actual product
Building the pipeline made this unavoidable. The model calls are a small fraction of it. The rest is the queue, the retries, the state, the storage, the credentials, the scheduling, the assembly. Script generation is one node. Getting from a spreadsheet row to an uploaded MP4 is a dozen more, and every one of them is a place where the whole thing can quietly stop working.
When I look at AI products now, I read the orchestration first: what happens between the calls, what state is being kept, what the system does with a half-finished job. That is where the product lives. The model is a component, an important one, and increasingly a commodity one.
What this is actually worth
I am not arguing that every designer should build a server. Most should not; it is a large amount of time to spend on something adjacent to the work.
What I am arguing is that operating a system produces a category of intuition you cannot get from using one, and that this particular gap is unusually wide right now. AI features are being designed at enormous scale by people whose entire mental model comes from a chat interface, and it shows, in the spinners, in the single-state errors, in the confident assumption that the model is the hard part.
If you want a cheaper version of the lesson: take one AI feature you have designed and try to build it end to end, badly, including the parts you would normally consider infrastructure. Not the model call. That is the easy bit, and it will work first try. The queue, the retry, the storage, the thing that runs when you are asleep.
What breaks is the curriculum.