We talk about Tanzu, but what are the differences between a supervisor cluster and a TKG cluster?
A Supervisor Cluster and a Tanzu Kubernetes Grid (TKG) Cluster have different roles in a VMware Kubernetes deployment.
VMware Supervisor Cluster

- Definition:
A Supervisor Cluster is a Kubernetes cluster that runs directly on vSphere using ESXi as the worker nodes.
It integrates Kubernetes natively with vSphere through vSphere with Tanzu.
Architecture:
Runs natively on ESXi, with each ESXi host serving as a Kubernetes worker node.
Uses vSphere Distributed Switch (vDS) and NSX-T for networking.
Incorporates VMware vSphere Pod Service, allowing native Kubernetes Pods to run alongside VMs.
Features:
vSphere Pods: Lightweight pods that run directly on ESXi, providing isolation and security similar to VMs.
Namespaces: Provide logical and security boundaries for resources within a vSphere environment.
Integrated Management: Manage through vCenter with vSphere roles and permissions.
Use Case:
Use it when Kubernetes workloads need direct integration with vSphere.
Tanzu Kubernetes Grid (TKG) Cluster

- Definition:
TKG clusters are Kubernetes clusters managed and deployed by VMware Tanzu Kubernetes Grid.
Can be deployed on multiple environments: vSphere, public clouds, and at the edge.
Architecture:
TKG clusters run on top of the Supervisor Cluster, but also support standalone deployments.
Deploys and manages clusters via Cluster API and Kubernetes Operators.
Features:
Multi-Cloud Support: Deploys across multiple cloud platforms like AWS, Azure, and vSphere.
Cluster API: Automates lifecycle management (creation, scaling, upgrade, and deletion) using Kubernetes-style declarative APIs.
Compatible: Works with standard Kubernetes tooling.
Use Case:
Use it to run consistent Kubernetes clusters across environments such as on-premises vSphere and public cloud.
Key Differences
- Deployment Model:
Supervisor Cluster: Kubernetes control plane runs directly on ESXi hosts.
TKG Cluster: Kubernetes clusters deployed on top of the Supervisor Cluster or on other platforms.
Network Integration:
Supervisor Cluster: Integrates deeply with NSX-T and vDS for networking and security.
TKG Cluster: Uses Calico for networking (NSX-T available for vSphere deployments).
Management Interface:
Supervisor Cluster: Managed via vCenter.
TKG Cluster: Managed via kubectl, Tanzu CLI, or through vCenter if deployed on a Supervisor Cluster.
Workload Types:
Supervisor Cluster: Supports vSphere Pods and Tanzu Kubernetes Clusters.
- TKG Cluster: Standard Kubernetes clusters for portable workloads.
In short
Supervisor Cluster provides native integration with vSphere and enables Kubernetes workloads to run directly on ESXi.
TKG Cluster offers consistent Kubernetes clusters across multiple environments.
TANZU Network choices:
So there is a network difference. This could be important in our design.
What separates the different network options:
Networking Solutions in TKG Clusters
- Calico:
Default Network Provider: In most TKG clusters, Calico is used as the default Container Network Interface (CNI).
- Features:
Network Policy: Implements Kubernetes NetworkPolicy for fine-grained traffic control.
IP Address Management: Manages pod IP addresses dynamically.
Overlay Networking: Uses VXLAN or IP-in-IP encapsulation.
Antrea:
Alternative Network Provider: In certain TKG clusters, Antrea is available as an alternative CNI.
- Features:
Open vSwitch (OVS) based networking.
Implements Kubernetes NetworkPolicy.
NSX-T Integration:
Available for vSphere Deployments:
When TKG clusters are deployed on vSphere with Tanzu (within Supervisor Clusters), NSX-T can be used as the network provider.
- Features:
Networking and Security Policies: Provides centralized network security policies via NSX-T.
Load Balancer: Offers built-in load balancing.
Networking: Supports Tier-0 and Tier-1 routing.
Choosing a Networking Solution
- Calico:

Use it when the environment needs a basic CNI.
Supports a broad range of TKG deployments.
Antrea:

Suitable for users looking for an OVS-based solution.
Provides OVS-based networking in TKG clusters.
NSX-T:

Use it when the environment needs NSX-T networking and security policies.
Integrates with vSphere.
Clarified Overview
- Supervisor Cluster (NSX-T or vDS):
NSX-T and vDS are used to provide networking.
Supervisor Cluster networks the TKG clusters deployed on top of it.
TKG Cluster:
Default CNI: Uses Calico or Antrea by default.
- NSX-T Integration: Available only for TKG clusters on vSphere.
NSX-T limitation
- Clarification:
NSX-T is not directly available in standalone TKG deployments but requires vSphere Supervisor Clusters.
So why would we have several TKG clusters in a single supervisor cluster?
1. Multi-Tenancy
Isolated Environments: Each TKG cluster can be allocated to different teams, departments, or tenants, ensuring that their resources, configurations, and security policies are isolated.
Access Control: Kubernetes RBAC can be applied independently within each TKG cluster, simplifying access management.
2. Workload Segmentation
Application Isolation: Different applications or microservices can be deployed in separate TKG clusters to minimize resource competition and security risks.
Environment Segregation:
Dev/Test/Prod Environments: Keep development, testing, and production workloads separate to avoid cross-environment issues.
- Compliance: Ensure compliance by separating applications that require different security policies or standards.
3. Resource Management
- Scalability:
Each TKG cluster can scale independently.
Allocates cluster resources according to workload needs.
Resource Quotas:
Each TKG cluster can be configured with quotas for CPU, memory, storage, etc.
- Prevents one team or tenant from monopolizing resources.
4. Application Modernization
- Legacy and Modern Applications:
Legacy applications requiring more control can be hosted in dedicated TKG clusters.
- Modern, cloud-native applications can be placed in separate clusters with different networking or security requirements.
5. Networking Customization
- Network Policies:
Different clusters can implement network policies using Calico or Antrea tailored to their specific requirements.
- NSX-T Integration:
NSX-T policies can provide centralized networking and security policies for each TKG cluster.
6. Disaster Recovery and High Availability
- Fault Isolation:
Multiple TKG clusters within the Supervisor Cluster minimize the impact of a single cluster failure.
- Backup and Restore:
Backup strategies can be specific to individual TKG clusters.
7. Lifecycle Management
- Rolling Updates:
Simplifies rolling updates since each TKG cluster can be updated independently.
- Cluster API (CAPI):
Cluster API manages the lifecycle of multiple TKG clusters.
Reasons to use several clusters
Multiple TKG clusters can separate teams, applications, lifecycle schedules, and security policies while sharing one Supervisor Cluster.