Skip to main content
Back to posts

🏠 My homelab has users: my parents

I wanted more say in where my family’s digital life runs.

At some point I came across the idea of setting aside a recurring day to move away from big tech services. Digital Independence Day does that on the first Sunday of each month: make one small, concrete change toward more independence. I liked that framing. For me it meant looking for European alternatives to US services.

I also wanted an AI family plan where I do the setup and my parents just use it. They should be able to ask a question without first researching models, providers and subscriptions, and I had not found an offering that worked like that.

So I built it myself. It now runs on seven Kubernetes nodes, which tells you the family plan was not my only interest.

A family plan, with my own plumbing

The experience I want is simple: open a page, sign in, ask for help with a manual or with planning dinner. Which inference provider answers should be my problem, not theirs.

Open WebUI is that page. It has German model labels like “Alltag” and “Schwierige Fragen”, speech input, Kokoro text-to-speech, and a family group for sharing useful material.

Behind it, LiteLLM exposes a configured set of models through one API. An upstream proxy sends requests to Eden AI, falls back to OpenRouter when needed, and notifies me when that happens. I can swap the provider behind the interface while my parents keep the familiar page.

The models themselves run at external providers. Hosting the interface gives me control over that layer. Prompts still leave the cluster for inference, and the fallback path can end up in a different region.

The lab runs a few other applications too:

ApplicationWhat it does here
TREKFamily travel planning
AFFiNECollaborative documents and planning
NocoDBSmall data applications
CloudBeaverPrivate database administration
Grafana and PyrraOperator dashboards and a family status view

Authentik provides single sign-on. Open WebUI, TREK and AFFiNE use it through OpenID Connect, as do several admin tools; the rest sit behind proxy authentication. Having identity in one place lets me keep the family and administrator roles apart.

European providers where I could

The provider choices are part of the design. For the infrastructure, the primary AI gateway, private networking and DNS, I picked companies based in the EU. Three of the four are German:

ProviderWhat I use it forCompany base
Eden AIPrimary gateway to external AI modelsLyon, France
HetznerCloud servers, volumes, and object storageGunzenhausen, Germany
NetBirdPrivate networking and managed network controlBerlin, Germany
INWXDomain and DNS managementBerlin, Germany

The wider setup still includes GitHub, 1Password and an OpenRouter fallback, so this is a work in progress. A provider’s headquarters also does not decide where every downstream AI request is processed.

What the architecture looks like

The machines run in Hetzner Cloud. Three Talos nodes form the Kubernetes control plane, four workers run the applications and platform services, and a separate NixOS utility host handles web routing and private network access.

Follow a family chat request or a visit to my private Grafana dashboard in the diagram. The numbered stops show where each request goes and when it leaves the lab:

THE LAB AT A GLANCE · SEPTEMBER 2026

One platform, two ways in

Family browserPublic app URLs
My devicesNetBird private network
HETZNER CLOUD
1 UTILITY HOST · NIXOS Caddy + NetBird routing peer Public app routes · private administration routes
Web traffic via explicit NodePorts
Talos · KubernetesCilium networking
3 control-plane nodes4 worker nodes
AuthentikSSO for integrated apps and admin tools
Family applicationsOpen WebUI · TREK
AFFiNE · NocoDB
Project workflowsSite build checks
Benchmark jobs & archives
Shared servicesPostgreSQL · 1Password Connect
Grafana · Prometheus · Loki · Velero
From the cluster: data & backups
HETZNER STORAGEVolumes + Object StorageLive data · backups · Pulumi state
From the cluster: AI via LiteLLM & proxy
EXTERNAL MODEL PROVIDERSEden AI → OpenRouter fallbackLanguage-model inference leaves the lab
DECLARED IN GIT
Pulumi / Python
Infrastructure + bootstrap
NixOS
Utility-host configuration
Argo CD
Kubernetes desired state
KDL → INWX
DNS zone via sync tasks
A map of responsibilities and web access, with selected services shown. Authentik is a shared identity service; apps integrate through OIDC or proxy authentication. NetBird also routes private API access. Storage and AI are separate dependencies.

Seven Kubernetes nodes for three people needs an explanation. Learning Kubernetes, GitOps and infrastructure automation is part of what I want from this project. The family services give that learning a purpose, and my other projects run on the same platform. The extra machinery is part of the hobby.

If you are building something similar, the question is which parts serve your own goals. A family chat page does not require copying my cluster. The choices around identity, provider access and reproducible configuration are useful on their own.

I chose Talos because it contains very little beyond what Kubernetes needs. No SSH daemon, no interactive shell, no package manager. Fewer components mean less to maintain and a smaller attack surface. Configuration and updates go through its API, which fits how I want to automate the lab.

The dedicated control plane came from experience. On the earlier mixed nodes, application spikes interfered with Kubernetes itself. Separating the roles helped, though the latest audit still found control-plane memory headroom to improve.

Caddy on the utility host routes web requests to explicit Kubernetes NodePorts, and Cilium handles cluster networking. That gives me a request path I can follow when something fails. The utility host is still a single point of failure, so its rebuild procedure matters too.

Everything is in Git

Pulumi, written in Python, provisions the Hetzner infrastructure and handles bootstrap work. Argo CD reconciles Kubernetes manifests and Helm values from Git. Infrastructure code lives in infra/, applications in cluster/, operating procedures in docs/. Invoke tasks provide the repeatable commands.

1Password holds the secrets; Connect and the Kubernetes operator deliver them to applications. Pulumi state lives in object storage. Dagger runs the CI checks, and Renovate proposes dependency updates.

The point of all this is that the intended state of the lab exists somewhere other than my memory.

The eighth machine gets NixOS

The utility host would be the easy place for that idea to fall apart. It sits outside Kubernetes and needs Caddy, NetBird, firewall rules and phone notifications through ntfy. Plenty of chances for a quick manual fix that becomes permanent because I forgot about it.

NixOS lets me describe that machine too. These are selected settings from its host module:

{ ... }:
{
services.caddy.enable = true;
services.netbird.enable = true;
services.netbird.useRoutingFeatures = "server";
services.openssh.enable = true;
services.openssh.settings.PasswordAuthentication = false;
}

The complete configuration includes the proxy routes and firewall rules. A flake assembles the system, its lock file pins the inputs, and Disko describes the disk layout. Python renderers fill in host-specific settings such as cluster addresses, so that wiring stays connected to the infrastructure tooling.

The machine starts from a temporary Linux image provisioned by Pulumi. My utility.install task uses nixos-anywhere to install the declared NixOS system. Later, utility.apply renders the configuration and runs a remote nixos-rebuild switch. Secrets come from 1Password; NetBird enrollment happens separately so its setup key stays out of Git.

One detail I especially like: Caddy has the INWX plugin and uses DNS challenges for certificates. It proves control of the domain through DNS, so a private dashboard can have browser-trusted HTTPS while only being reachable over the private network. NixOS, INWX and NetBird each do one part of that.

Even DNS gets a file

The DNS records live in a KDL file, with INWX as the provider.

Here is a shortened example in the same format as my real configuration. Domain and IP are documentation examples:

zone "example.org" {
provider "inwx"
record "chat" type="A" value="192.0.2.10" ttl=300
record "auth" type="A" value="192.0.2.10" ttl=300
record "travel" type="A" value="192.0.2.10" ttl=300
}

One zone, one provider, and records with type, value and time to live. The real file also has the mail and verification records. They belong to the domain as much as the homelab’s web addresses do.

Repository tasks parse the file, compare it with the live INWX zone, and show which records would be created, updated or deleted. DNS uses this explicit sync workflow next to Pulumi:

Terminal window
uv run invoke dns.plan
# Review the proposed changes, then apply them.
uv run invoke dns.apply
uv run invoke dns.verify --server ns.inwx.de

There are checks for nameserver delegation and propagation too. For a provider migration, the sequence is: inventory the old zone, declare it in Git, populate and verify the new provider, then change delegation. A DNS change is a diff I can read before applying it.

NetBird for the private network

Along the way I found NetBird, an alternative to Tailscale from a Berlin-based company. Its clients and core server components are open source, including the management, signal and relay services. A self-hostable option appealed to the same instinct that started the lab.

I currently use NetBird Cloud for management. The utility host is a routing peer into the private network, which gives my devices access to cluster administration and private tools like Argo CD, Grafana and CloudBeaver. My parents use the public family apps without joining that network.

A home for my other projects

The cluster also runs the work behind my projects.

Argo Workflows refreshes this website’s data, builds it, runs the browser checks and publishes the tested result every three hours. The public site stays on Netlify; the PostgreSQL data behind it stays private in the cluster.

My speed-comparison project has its benchmark workflows and a private result archive here too. The benchmark workflow pins a specific worker for a consistent machine, though neighboring workloads can still add noise. The weekly benchmark schedule is suspended while that workflow evolves; a separate daily job imports published snapshots into the archive.

These projects share the scheduling, storage, secrets and monitoring that already exist for the family services. That makes the lab a small development platform as well.

Backups, and a failed restore

Sharing infrastructure also means sharing limits. Several applications use a three-instance PostgreSQL cluster managed by CloudNativePG, with separate databases and roles. After a burst of login traffic used up its connection budget, I capped Authentik’s role to leave headroom for the neighbors.

The most useful backup result so far was a failed restore.

PostgreSQL has database-aware backups and WAL archiving; Velero covers Kubernetes resources and the relevant filesystem data. During a drill in September, PostgreSQL came back but the CloudBeaver filesystem restore never started. The verification step caught it. After fixing the resource filter, the drill restored both into fresh volumes and checked their contents.

That is worth more than a green backup indicator: a mistake found while the original data was still there. It was a component drill on a working cluster, and a full rebuild still needs its own exercise.

Prometheus, Grafana, Loki and phone notifications help me run it, with alerts focused on sustained failures. All of it serves the original goal: a service for my parents that I can understand, change and recover.

I get somewhere to learn and to run my projects. My parents get a page in the browser. Ideally they never think about the seven nodes behind it.