Skip to Content
Architecture Overview

Architecture Overview

kubenest is a distributed system with five cooperating components. Each component has a single, well-bounded responsibility, and the boundaries between them are deliberately strict. This page explains what each component does, how they communicate, and why the system is designed the way it is.

Understanding the architecture pays off quickly: it tells you where to look when a deployment stalls, why certain operations are async, and exactly what the control plane can and cannot reach inside your cluster.

The five components

kubenest-backend

The backend is the control plane. It is a FastAPI application that owns every user-facing resource: organizations, users, clusters, projects, apps, stack templates, addon instances, and deployment history. All persistent state lives in PostgreSQL. Redis handles job queues, rate-limit counters, and short-lived cache entries (including chart inspection results).

Application, addon and component-secret operations are dispatched as schema-validated JSON events to the hub and acknowledged asynchronously, rather than executed against your cluster from inside the backend. The control plane sends intents through the hub and reads what happened from the status events the operators send back. No request state is held in memory between calls, so you can run multiple replicas behind a load balancer without any shared in-process state.

kubenest-hub

The hub is a Go WebSocket message broker. Its sole job is routing events between the backend and the operators running on each cluster. The backend connects to the hub as a persistent WebSocket client (identified by a client_type: "backend" claim in its JWT). Each operator also connects as a persistent client (authenticated by a cluster-scoped JWT). The hub maintains a connection registry in Redis so that multiple hub replicas can be deployed without routing failures.

The hub understands two event directions: command events flow backend → operator (deploy, patch, delete, redeploy); status events flow operator → backend (phase transitions, health checks, drift reports, log lines). The hub does not inspect or transform event payloads — it validates the sender’s identity and routes the message to the correct recipient session.

kubenest-operator

The operator is a Go controller-runtime application that runs inside each registered Kubernetes cluster. It implements five Kubernetes controllers:

  • StackDeploy controller — watches StackDeploy CRDs (the cluster-side representation of an App) and drives the full deploy/patch/delete reconcile loop.
  • Workload controller — tracks individual workload components within a StackDeploy: watches the Argo CD Application’s health, reports the cluster’s ingress IP, and collects exports once a component is running. The operator creates that Application on its own cluster. See GitOps.
  • Addon controller — same as Workload but for Helm-chart backing services; also handles export discovery after a successful install.
  • BuildRequest controller — manages image build jobs (Kaniko or Cloud Native Buildpacks) for Dockerfile and buildpack deployment modes.
  • Project controller — ensures the Kubernetes namespace exists and synchronizes an optional project registry secret before any app is deployed into it. It does not manage resource quotas or RBAC bindings.

There is exactly one operator instance per cluster. The operator has local Kubernetes API access and uses that access exclusively — it never routes Kubernetes calls through the hub or the backend.

kubenest-ui

The web console is a Next.js application. It communicates with the backend over REST for all CRUD operations (creating apps, listing clusters, reading deployment history). Real-time status — phase changes, ArgoCD sync progress, log streams — arrives via a Server-Sent Events subscription to the backend’s /api/v1/events/stream endpoint. The UI never connects directly to the hub or to the operator; it sees only what the backend surfaces.

kubenest-contracts

The contracts repository is not a running service — it is the shared source of truth for the JSON Schema definitions of every event that flows through the system. Command events and status events are validated against these schemas at the hub boundary: the hub rejects any malformed event before routing it. This guarantees that the operator never receives a structurally invalid command, and the backend never receives a structurally invalid status update.


System topology

The backend and each in-cluster operator open outbound authenticated WebSocket connections to the hub; only the operator connects to the local Kubernetes API

The data flow for the two main directions:


Three core design principles

1. Cluster access

Most Kubernetes work originates inside the target cluster. The agent holds the local API access and acts on schema-validated events routed through the hub, and for application, addon and component-secret operations the control plane builds no Kubernetes client for your cluster at all.

Every workload Argo CD Application is created by the operator on its own cluster. The backend records the intent and dispatches it as an event; the operator on the target cluster turns that intent into the Application and reports status back through the hub. A stack’s components all run on the stack’s cluster.

The agent’s own session requires no inbound firewall rule. The operator dials the hub outbound over WebSocket and holds the connection open, so it survives NAT and a restrictive corporate firewall as long as outbound HTTPS/WSS is allowed.

2. ArgoCD reconciles everything; Git backs the addons

Nothing in kubenest applies Kubernetes resources with kubectl apply. ArgoCD does the applying, in both paths. Where the two paths differ is where the desired state lives, and the difference is worth knowing before you go looking for a commit that was never made.

Addons are backed by Git. The addon controller writes the chart and values to a per-cluster GitOps repository under clusters/{id}/addons/, then points an ArgoCD Application at that directory. An addon therefore has a human-readable representation in Git that you can inspect, diff and revert independently of kubenest, and ArgoCD keeps serving it if kubenest is offline.

Workloads are not. The operator creates the Argo CD Application on the cluster it runs in, with the Helm values inline in the Application spec against a chart repository. No workload manifest is committed, and there is no workload directory in the GitOps repo. A workload’s change history lives in the backend’s deployment records rather than in Git.

GitOps has the full comparison, including what this costs you.

3. Event-driven asynchrony for all cluster operations

Cluster operations are long-running and unreliable — network partitions, slow image pulls, failing readiness probes. kubenest models all of them as asynchronous events rather than synchronous RPC. When you POST an App create request, the backend records the intent and immediately returns a Pending response. The operator reconciles on its own schedule and emits status events as it makes progress. The UI reflects those events in real time through the SSE stream.

This architecture eliminates HTTP timeouts from the critical path and makes the system resilient to transient connectivity failures between the backend and the hub, or between the hub and the operator.

Command delivery guarantee

The hub is a router, not a durable command queue. A backend write that needs an operator command must use one of two contracts:

  • Durable intent. Persist the complete desired state before dispatching, return Pending or Accepted, and keep deriving the command until a matching operator observation confirms the cluster state. If confirmation is delayed, the reader exposes a named degraded lifecycle reason while reconciliation continues.
  • Fail closed. If an operation has no durable state that can be reconciled, dispatch without relying on the client-side queue and return a non-2xx response when the hub cannot confirm the route.

A socket write, a connected preflight, or admission to a temporary client queue never proves that a cluster changed. A hub route confirmation only proves routing; the operator’s status observation is what confirms cluster state.


Request flow patterns

Synchronous operations

Some operations complete entirely within the backend and return a result immediately. These are pure metadata operations that do not touch Kubernetes.

Examples: creating a project record, listing clusters, reading deployment history, fetching a stack template.

Asynchronous cluster operations

Operations that change Kubernetes state are asynchronous. The backend dispatches a command event and immediately returns; the operator reconciles and streams status back.

Status update loop

The operator runs a continuous reconcile loop. Whenever ArgoCD reports a sync or health change, the operator emits a status event that propagates all the way to the UI:


Security architecture

User authentication

Users authenticate with a username/password POST to /api/v1/login. The backend issues two tokens:

  • Access token — a short-lived JWT (30-minute expiry) sent in the response body. Clients include it in the Authorization: Bearer header on every request.
  • Refresh token — a 7-day JWT stored in an HttpOnly, Secure, SameSite=Strict cookie. Clients call POST /api/v1/refresh to get a new access token when the current one expires. Because the refresh token is in an HttpOnly cookie, JavaScript cannot read it, mitigating XSS-based token theft.

Access tokens are validated on every request by the backend’s FastAPI dependency injection layer. There is no token introspection call to an external service — the backend validates the signature locally using the shared secret.

Operator authentication

When a cluster is registered, the backend mints a cluster JWT — a long-lived token scoped to that cluster’s UUID. This token is embedded in the Kubernetes Secret created by the operator Helm chart. The operator presents it to the hub in the Authorization: Bearer header when establishing its WebSocket connection.

The hub validates the cluster JWT (signature + cluster_id claim), then checks the token’s token_version claim against the cluster’s revocation floor — the lowest version the hub still admits. A token below the floor is refused, and a rotation also closes the cluster’s live session rather than waiting for it to reconnect. Once admitted, the session is associated with that cluster ID and subsequent events from it are authoritative for that cluster — the backend trusts hub-routed status events without re-validating the operator’s identity on each message.

The floors are pushed to the hub by the backend and held in the hub’s memory. A hub that has restarted and has not yet received them, or whose copy has gone stale, refuses every operator connection with a retryable 503 until the control plane reconnects. That is deliberate: a hub restart is predictable, so admitting operators unverified for a window after one would make the window selectable. Existing connections are not dropped and workloads on the cluster are unaffected — this is a control-plane pause, not an outage on your cluster.

The cluster JWT is long-lived and has significant privilege. Treat it like an SSH private key: store it in a Kubernetes Secret (the Helm chart does this automatically), restrict access to that Secret via RBAC, and revoke it immediately if you believe it was compromised.

Revoke with POST /clusters/{id}/rotate-token — see the API reference. It is the only call that raises the revocation floor, and every token previously issued for that cluster stops working the moment it returns.

It is not a hygiene operation. The new token reaches the agent through the operator chart’s values, so the cluster stays disconnected until you run an install or upgrade carrying it. Rotating a fleet on a schedule would disconnect the fleet. Re-running kubenest platform install is not a revocation: it issues a higher-versioned token and deliberately leaves the floor where it is, so re-running an install cannot lock a healthy cluster out of its own hub.

Backend-to-hub authentication

The backend authenticates to the hub using a JWT with a client_type: "backend" claim. This is distinct from cluster JWTs and from user access tokens. The hub uses the client_type claim to apply different routing rules: messages from the backend session are forwarded to operator sessions, and vice versa.

Network boundaries

ConnectionProtocolAuthentication
UI → BackendHTTPS REST / SSEUser access token (Bearer)
Backend → HubWSS (persistent)Backend JWT (client_type: backend)
Operator → HubWSS (persistent)Cluster JWT
Operator → Kubernetes APIIn-cluster HTTPServiceAccount token (RBAC-scoped)
Operator → GitHTTPSPersonal access token (read/write on GitOps repo only)
ArgoCD → GitHTTPSSame GitOps PAT
ArgoCD → Kubernetes APIIn-clusterArgoCD ServiceAccount

Multi-tenancy model

kubenest enforces isolation at four levels, each reinforcing the others.

Organization level (backend). Every database row carrying user data has an org_id foreign key. All API queries include WHERE org_id = <current_org> — organization A cannot read organization B’s data regardless of how the query is constructed. Backend middleware enforces this before any route handler runs.

Cluster level (routing). A cluster’s operator session in the hub is keyed by cluster_id. The backend can only dispatch events to clusters that belong to the authenticated user’s organization. The hub enforces this routing — you cannot send a command event to a cluster you do not own by constructing a crafted WebSocket message, because the hub validates the sender’s org_id claim against the target cluster’s registration.

Project level (Kubernetes). Each project maps to a Kubernetes namespace. The operator creates that namespace before deploying any application into it. Namespace scoping keeps objects in project A separate from project B; KubeNest does not currently manage resource quotas or RBAC bindings for those namespaces.

Namespace boundary (operator). The operator’s Project controller enforces namespace boundaries on all reconcile operations. A StackDeploy that specifies a target namespace outside its registered project is rejected by the operator with a validation error that propagates back through the hub to the backend.

GitOps

ArgoCD is the common reconciliation layer. Git is the source of truth for addons, and not for workloads — a distinction worth knowing before you go looking for something in the wrong place.

Workloads are created by the operator as Argo CD Applications on the cluster it runs in, with inline Helm values and no Git commit, and history in PostgreSQL. Addons are reconciled by the in-cluster addon controller, which commits Chart.yaml and values.yaml under the cluster's addon directory and points an Argo CD Application at that Git path.

Workload history lives in PostgreSQL. Every App operation is a Deployment row, and the ones that support rollback store the prior StackDeploy snapshot alongside. Workload ArgoCD Applications carry inline Helm values, so a git log is not the audit trail for a workload deploy, patch or rollback — the API is.

Addon history lives in Git. Each addon reconciliation writes Chart.yaml and values.yaml under the cluster’s addon directory and commits:

gitops-repo/ └── clusters/ └── {cluster-id}/ └── addons/ └── {namespace}-{addon-name}/ ├── Chart.yaml └── values.yaml

Either way ArgoCD self-heals against its desired state, prunes what should not be there, and reports sync and health back.

Drift detection

Drift is live Kubernetes resources diverging from desired state. ArgoCD reports Application sync and health; the Workload controller separately compares the desired image and replica count against the observed Deployment, and the StackDeploy controller aggregates its children.

The API exposes it as sync on the App:

{ "phase": "Running", "sync": { "driftDetected": true, "driftClass": "recoverable", "driftDetails": [ { "resource": "Deployment/my-app-web", "field": "spec.replicas", "desired": 2, "observed": 1 } ] } }

Two classes, distinguished by what they do to your next operation:

  • recoverable — ArgoCD can re-apply the desired resources itself, and does.
  • blocked_sync — the Application is Unknown, Error or Degraded. Mutating routes return 409 until the underlying condition is resolved, rather than layering another change on top of a cluster that is not converging.

Treat workload ArgoCD Applications and addon Git directories as KubeNest-managed desired state. Direct edits can be overwritten by the next API operation or controller reconciliation.


See also:

Last updated on