I deployed VMware Cloud Foundation (VCF) 9.1.1 on three nested ESXi hosts running on a single Minisforum MS-A2, then used vSphere Kubernetes Service (VKS) to run NGINX and access it from my Mac. This walkthrough covers the deployment, networking, storage validation, and fixes along the way.
William Lam’s work provided the foundation. I followed his TEP-less VLAN-backed VPC walkthrough and nested VCF deployment guide, adapting his upstream automation scripts for my Ryzen hardware and network.
1. Understand the design and prepare the lab #
The lab uses a VLAN-backed VPC (Virtual Private Cloud) for workload networking. A DTGW (Distributed Transit Gateway) connects it to the external network, and a VNA (Virtual Network Appliance) provides network services. Both remain part of the TEP-less design.
The Supervisor provides the vSphere-integrated Kubernetes control plane. VKS provisions the guest Kubernetes cluster that runs the applications. Its nodes use subnets allocated from a Public SubnetSet in the namespace.
TEP-less means the ESXi hosts do not need Tunnel Endpoint IPs; their TEP assignment was NO_IP. The guest cluster uses Antrea for pod networking. Antrea supports multiple traffic modes, and I did not verify which was active, so TEP-less does not imply overlay-free pod networking.
The diagram shows logical relationships; application traffic does not pass through the Supervisor. The outer vCenter manages the nested infrastructure; the nested vCenter manages the VCF domain.
| Foundation | My recorded configuration |
|---|---|
| Physical computer | One Minisforum MS-A2; AMD Ryzen 9 9955HX |
| Physical memory | 128 GB DRAM; NVMe Memory Tiering already configured |
| Router | MikroTik RB5009 |
| Outer vCenter | vc01.lab.devynharrington.com |
| Outer cluster / datastore | Lab-Cluster / nested-vcf |
| Nested hosts | nested-esx01, nested-esx02, nested-esx03 |
| Host management | 192.168.88.41, .42, .43; /24 |
| Per-host configured resources | 24 vCPU, 128 GB memory |
| Recorded per-host virtual disks | 32 GB cache VMDK; 3000 GB capacity VMDK |
| Installer | inst01.vcf.lab.devynharrington.com; 192.168.88.50 |
| Management gateway / DNS | 192.168.88.1 |
| DNS suffix / NTP | vcf.lab.devynharrington.com / time.google.com |
| Nested vCenter | vc01.vcf.lab.devynharrington.com |
| Nested datacenter / cluster | vcf-mgmt-dc / vcf-mgmt-cl01 |
| External workload network | VLAN 2006; 10.1.35.0/24; gateway 10.1.35.1 |
The router interface vlan2006-external uses 10.1.35.1/24. Use the same gateway CIDR for the transit gateway’s external connection and 10.1.35.0/24 for the External IP Block.
The outer inventory reported approximately 628 GB of tiered capacity, not physical DRAM. All three nested hosts share the physical host’s CPU, memory, storage, and failure domain. The cache/capacity labels above come from the deployment script, not general vSAN ESA sizing guidance.
My earlier physical-host and VCF article covers the existing foundation. Its six-host topology, Automation deployment, and JSON bring-up are a different build. Here I retained the physical host, outer vCenter, datastore, router, and local assets, then rebuilt the three-host nested domain.
Prerequisites to check before starting #
The upstream repository calls for an outer vCenter-managed environment, PowerShell Core/PowerCLI, the nested ESXi and Installer OVAs, an entitled software source, and networking that permits nested guests.
| Requirement | What to prepare or verify |
|---|---|
| Software access | Entitlement to the VCF binaries and an online or offline depot. For 9.1, the documented online workflow registers a Software Depot ID in VCF Business Services and returns an Activation Code. |
| Local tools | PowerShell (pwsh) and compatible PowerCLI; vcf and kubectl for later access; SSH for RouterOS. |
| Nested virtualization | Outer ESXi must expose hardware virtualization to the nested hosts and have enough storage and memory to run the appliances. |
| Outer networking | A management path and a trunk path allowing the required nested MAC addresses and VLANs. Use MAC Learning or promiscuous mode as appropriate to the outer switch, and enable forged transmits with either choice. |
| DNS and time | Unique appliance FQDNs, working forward/reverse DNS as required by the Installer, reachable gateways and NTP, and client access to both management and external networks. |
| Storage | Space on nested-vcf for the nested VMDKs and on nested vsanDatastore for management appliances, libraries, and guest VMs. Thin provisioned capacity is not a reservation of physical free space. |
| Supervisor readiness | vSphere HA enabled and relevant cluster/vSAN health prerequisites checked before activation. |
The Broadcom PowerCLI guide recommends PowerShell 7.4 or later and the VCF.PowerCLI module. Review the VCF 9.1 depot activation article and Cloud Proxy DNS prerequisite article before bring-up.
Use your own domain, IPs, datastore and cluster names, port groups, and local paths. Generated names and allocated addresses will vary by deployment.
2. Deploy the nested hosts and Installer from PowerShell #
I worked in ~/VCF91-Lab on the Mac with the following scripts and OVAs:
| Asset | Local filename |
|---|---|
| Customized deployment driver | vcf-automated-fleet-deployment-9.1.1-tepless.ps1 |
| Customized environment settings | devyn-vcf-9.1.1-tepless-no-automation-3tb.ps1 |
| Supporting local test script | Test-DevynVCF911TeplessJson.ps1 |
| Nested ESXi OVA | Nested_ESXi9.1.0.0_Appliance_Template_v1.0.ova |
| Installer appliance OVA | VCF-SDDC-Manager-Appliance-9.1.1.0.25713928.ova |
Start with William’s repository and adapt its sample configuration to your environment. The filenames and -Phase Infrastructure parameter below belong to my customized scripts.
Mac terminal — zsh, then enter PowerShell:
cd ~/VCF91-Lab
pwshPowerShell on the Mac — recorded local Infrastructure phase:
./vcf-automated-fleet-deployment-9.1.1-tepless.ps1 `
-EnvConfigFile ./devyn-vcf-9.1.1-tepless-no-automation-3tb.ps1 `
-Phase InfrastructureThe PowerShell backtick continues a line; do not put spaces after it. The Infrastructure phase prepared the nested hosts and Installer. The subsequent management-domain deployment was performed through the Installer UI.
Before accepting the script’s Y/N prompt, I checked VCF 9.1.1.0, three hosts, 24 vCPU/128 GB each, 32 GB cache and 3000 GB capacity VMDKs, Day-0 Automation DISABLED, and TEP assignment NONE with overlayVtepSpec.vtepType = NO_IP. The summary also enabled the nested-lab vSAN ESA HCL auto-disk-claim exception. The Installer review later calls this Allow auto claim of HCL incompatible disks. It is a lab accommodation, not a statement of supported physical hardware.
I recreated the nested lab and reinstalled the 9.1.1 binaries before the successful rebuild.
Checkpoint: in the outer vCenter, verify the three nested hosts and inst01, including their networks and configured resources. Confirm access to https://inst01.vcf.lab.devynharrington.com and prepare the entitled 9.1.1 binaries before starting the fleet wizard.
3. Complete the successful VCF Installer UI workflow #
I opened the Installer, chose Deploy a new VCF fleet, and used the new-management-domain path without existing components.
In Plan, I selected the Simple deployment model and Small deployment size. The selection below shows the settings used for the successful rebuild.
Select VLAN-backed VPC in Network Options #
In Plan → Network Options, click Customize. Under VPC Network Configuration, change Full Stack VPC to VLAN backed VPC.
The Distributed and Centralized gateway choices disappear when VLAN backed VPC is selected. Confirm this selection before clicking Next. It sets the host TEP configuration to overlayVtepSpec.vtepType = NO_IP; the external VLAN, gateway, IP block, and VNA are configured after deployment in section 5.
General information and host entry #
In Prepare, I entered the following values. This table follows the final review, so it supersedes earlier default screens.
| Field | Successful value |
|---|---|
| Version | 9.1.1.0 |
| VCF instance | Devyn VCF 9.1.1 Lab - TEP-Less VLAN-Backed VPC |
| Management domain | vcf-m01 |
| CEIP | No |
| Default hostname DNS suffix | vcf.lab.devynharrington.com |
| DNS / NTP | 192.168.88.1 / time.google.com |
| Autogenerate passwords | Unchecked; chosen passwords entered manually |
| Day-0 Automation | Disabled |
| Hosts | nested-esx01.vcf.lab.devynharrington.com, nested-esx02.vcf.lab.devynharrington.com, nested-esx03.vcf.lab.devynharrington.com |
General Information form before final settings
For each host, enter its current root password and verify the presented certificate fingerprint against the host. Rebuilding a host creates a new certificate, so do not reuse a previous deployment’s thumbprint.
Networks and the management-services address pool #
Enter the ESX, VM, and VCF management networks, then the vMotion and vSAN networks. Keep the management-services pool separate from the external workload block.
| Network | VLAN | CIDR / gateway | Address range / MTU |
|---|---|---|---|
| ESX Management | 0 |
192.168.88.0/24 / 192.168.88.1 |
Existing host addresses; no separate MTU shown in review |
| VM Management | 0 |
192.168.88.0/24 / 192.168.88.1 |
Management appliance network |
| VCF Management | Inherit | Use same inputs from VM Management Network | Inherit |
| VCF Management Services pool | Inherit | Management subnet | 192.168.88.193–192.168.88.222 |
| vMotion | 0 |
10.1.32.0/24 / 10.1.32.1 |
10.1.32.101–10.1.32.118; MTU 9000 |
| vSAN | 0 |
10.1.33.0/24 / 10.1.33.1 |
10.1.33.101–10.1.33.118; MTU 9000 |
I selected Single IPv4 range for the 30-address management-services pool listed above; the form required at least 12. The DHCP pool ended at .191. The VNA and Supervisor floating addresses, .90 and .91, were outside both pools.
VLAN 0 applies to these nested networks. The external workload network uses VLAN 2006.
Operations, Cloud Proxy, and management services #
| UI field | Successful value |
|---|---|
| Operations size / primary node | Small / ops01.vcf.lab.devynharrington.com |
| Cloud Proxy size / FQDN | Small / vcf-proxy01.vcf.lab.devynharrington.com |
| License Server | vcf-lic01.vcf.lab.devynharrington.com |
| VCF services runtime | vcf-msr01.vcf.lab.devynharrington.com |
| Instance components | vcf-int01.vcf.lab.devynharrington.com |
| Fleet components | vcf-flt01.vcf.lab.devynharrington.com |
| Identity Broker | vcf-idb01.vcf.lab.devynharrington.com |
| Management-services size / name | Small / vcf-m01-vmsp-01 |
| Internal cluster CIDR IPv4 | 198.18.0.0/15 |
Enter the corresponding root, administrator, and system-user passwords in the wizard.
Nested vCenter, storage, and distributed switch #
| UI field | Successful value |
|---|---|
| vCenter | vc01.vcf.lab.devynharrington.com |
| Appliance size / appliance storage size | Small / Large |
| Datacenter / cluster | vcf-mgmt-dc / vcf-mgmt-cl01 |
| SSO domain | vsphere.local |
| Storage | vSAN ESA |
| Allow auto claim of HCL incompatible disks | Yes, for this nested lab |
| Datastore | vsanDatastore |
| Data-in-Transit encryption | No |
| VDS | vcf-mgmt-cl01-vds01 |
| VDS MTU / uplink type / count | 9000 / VDS Uplinks / 2 |
| Installer NIC mapping | vmnic0 → uplink1; vmnic1 → uplink2 |
I kept the Installer mapping on vmnic0/vmnic1. The trunk migration comes after deployment, as section 4 explains.
| Traffic | Port group | Load balancing | Active uplinks |
|---|---|---|---|
| ESX Management | vcf-mgmt-cl01-vds01-pg-esx-mgmt |
Route Based on Physical NIC Load | uplink1, uplink2 |
| VM Management | vcf-mgmt-cl01-vds01-pg-vm-mgmt |
Route Based on Physical NIC Load | uplink1, uplink2 |
| vMotion | vcf-mgmt-cl01-vds01-pg-vmotion |
Route Based on Physical NIC Load | uplink1, uplink2 |
| vSAN | vcf-mgmt-cl01-vds01-pg-vsan |
Route Based on Physical NIC Load | uplink1, uplink2 |
NSX, SDDC Manager, validation, and completion #
The VLAN backed VPC selection made in Plan → Network Options is retained in the saved configuration as vtepType: NO_IP.
| UI field | Successful value |
|---|---|
| NSX operational mode | Apply the default virtual switch mode configured in NSX Manager |
| Overlay transport-zone name retained in review | overlay-tz-mgmt-nsxt |
| NSX Manager size | Medium |
| NSX appliance / cluster FQDN | nsx01a.vcf.lab.devynharrington.com / nsx01.vcf.lab.devynharrington.com |
| NSX addresses observed later | Node 192.168.88.61; cluster VIP 192.168.88.60 |
| SDDC Manager | sddcm01.vcf.lab.devynharrington.com |
| VCF Installer Deployment | New SDDC Manager Deployment |
The overlay transport-zone name remains in the review even though host TEP IPs are not assigned. Enter the NSX root/admin/audit and SDDC Manager root/VCF/admin credentials, review the specification, and run validation. Resolve any errors before starting deployment. Keep exported specifications private because they can contain credentials.
The Installer reported a successful deployment:
About 5 hours 43 minutes elapsed between the deployment-start and completion screenshots in my nested lab. Confirm the green stage results and success banner before continuing to postdeployment networking.
4. Move the completed VDS onto the outer trunk adapters #
The Installer used vmnic0/vmnic1. After deployment, I moved the VDS uplinks to the outer-trunk-backed vmnic2/vmnic3 adapters for the VLAN-backed workload network.
My earlier attempt to specify vmnic2/vmnic3 too soon failed ESX Host Configuration validation: vmnic0 was attached to vSwitch0, but was missing from the input specification. That is why I retained vmnic0/vmnic1 during the successful Installer run.
In outer vCenter → VM → Edit Settings, I checked each nested host’s extra adapters were connected to VCF-Trunk, configured with VLAN 4095. Match adapter MAC addresses to the nested host’s physical-adapter inventory to confirm the mapping.
The Broadcom nested-VLAN reference distinguishes outer switch types: VLAN 4095 is the standard-port-group virtual guest tagging setting; a distributed port group instead uses VLAN Trunking with an allowed range. Permit VLAN 2006 across the actual path and retain the management/storage connectivity required by the nested hosts.
In the nested vCenter → Networking → vcf-mgmt-cl01-vds01 → Add and Manage Hosts, I selected the nested hosts and managed their physical adapters.
| Nested adapter | During Installer | Target after deployment |
|---|---|---|
vmnic0 |
uplink1, outer VM Network |
Unassigned from this VDS |
vmnic1 |
uplink2, outer VM Network |
Unassigned from this VDS |
vmnic2 |
Available, outer VCF-Trunk | uplink1 |
vmnic3 |
Available, outer VCF-Trunk | uplink2 |
Apply changes deliberately, preserving a verified working path while migrating a host. Review the wizard’s impact and keep console access available. Do not detach all working adapters at once or assign the same host’s uplink to two adapters. After each host, verify management reachability and storage networking before proceeding.
I completed the trunk correction during the rebuild. Figure S24 shows the starting state; Figure H181 illustrates the final mapping from a previous deployment.
5. Configure the transit gateway, external block, and Large VNA #
With VCF up and the trunk path corrected, I rechecked the router’s /24.
In nested vCenter → Network → Networks → Transit Gateways, click Set Up Network Connectivity to begin.
The prerequisite screen calls for the external VLAN path and disabling ICMP redirects on its gateway. I verified the intended VLAN connectivity, then applied the recorded RouterOS setting.
MikroTik RouterOS terminal — disable and verify ICMP redirects:
/ip/settings/set send-redirects=no
/ip/settings/printIn External Network Connectivity, enter VLAN 2006 and gateway CIDR 10.1.35.1/24. Add the block in the same workflow:
| Field | Submitted value |
|---|---|
| External IP Block | VCF-External-IP-Block |
| Visibility | External |
| CIDR | 10.1.35.0/24 |
| Additional IP ranges | Blank |
| Excluded range | 10.1.35.1–10.1.35.1 |
| Reserved for Specific Subnet | No |
| Description | VLAN 2006 external network for TEP-less VPCs |
External-connection and IP-block entry forms
The gateway field uses the host address, 10.1.35.1/24; the block uses the network, 10.1.35.0/24. Exclude the router address and any other reserved addresses from allocation.
The transit gateway provides the external connection; the VPC connectivity profile links the VPC to the external IP block and services. My Supervisor used the Default NSX project and Default VPC Connectivity Profile. Verify that profile references this /24 block before enabling Supervisor.
Submit the VNA configuration #
On VPC Services, configure the VNA cluster and add its node.
| VNA setting | Recorded value |
|---|---|
| Cluster / node | vcf-m01-vna-cl01 / vna01 |
| Form factor / node count | Large / one |
| Placement | vcf-mgmt-cl01; default cluster resource pool |
| Datastore | vsanDatastore |
| Host group affinity | No |
| Management port group | vcf-mgmt-cl01-vds01-pg-vm-mgmt |
| Management address | 192.168.88.90 on 192.168.88.0/24 |
| Gateway / DNS | 192.168.88.1 / 192.168.88.1 |
| Planned FQDN | vna01.vcf.lab.devynharrington.com |
Intermediate VNA forms: Medium and DHCP are initial defaults
The initial service form showed Medium, and the unfilled node form showed DHCP. Neither establishes the final submission. The recorded management plan used .90; the healthy result later confirms that address. The submitted review conclusively shows Large, one node, and the correct external connection:
Monitor deployment under Configure → Networking → VNA Clusters. Figure S34 shows vna01 at 10% Deploying VM.
6. Recover VNA root access and apply the Ryzen workaround #
The VNA did not deploy successfully on my Ryzen lab without intervention. Deployment stalled at around 51%, the same symptom I had encountered in the earlier attempt. The VNA hit an AMD CPU-model validation guard: the installed Python code rejected an AMD processor unless its model string contained AMD EPYC.
To get deployment moving again, I opened vna01’s VM web console in the nested vCenter, restarted the appliance, and used GRUB password recovery to reset its root password. Once I could log in as root, I applied the sed workaround below, verified the change, and rebooted. The recovery steps and commands are shown in order below; the stall near 51% was an observation from my lab.
William’s original Ryzen/NSX Edge article explains the guard and workaround. This is an unsupported homelab modification. His upgrade follow-up covers changes to file paths and line numbers; upgrades or redeployments can replace the modified file.
Use the VNA console and GRUB password recovery #
- In the nested vCenter, select
vna01, launch its VM web console, and restart the appliance. - Hold left Shift or press Esc during boot to interrupt GRUB. If you miss the menu, restart and try again. Select the boot entry and press e to edit.
- If GRUB authenticates the edit, use its credentials. The documented unchanged default is username
root, passwordNSX@VM!WaR10. A changed GRUB password takes precedence. This is not the appliance’s new root login password. - Locate the line beginning with
linuxand move to the end of the logical line, even if it wraps across the display. On the Mac, the guidance was Ctrl+E or Fn+Right; confirm the actual cursor position. - Append a space and the recovery parameter below.
- Press Ctrl+X to boot. Follow the recovery prompt to set and confirm the appliance root password, then complete boot and log in as root.
VNA VM console — append to the GRUB linux line, not a shell command:
systemd.wants=PasswordRecovery.serviceThe parameter and boot workflow are documented in Broadcom KB 416526 and the Edge/Manager recovery procedure in KB 316043. I reached the VNA root shell through this process. No recovered appliance password is included in this article.
Console message about the separate GRUB password
Inspect the installed file before editing #
VNA root shell — advised backup and recorded inspection command:
cp -pn /opt/vmware/nsx-edge/bin/config.py \
/opt/vmware/nsx-edge/bin/config.py.pre-ryzen-workaround
grep -n -A5 -B5 'Unsupported CPU\|AMD EPYC' \
/opt/vmware/nsx-edge/bin/config.pyThe backup was advised in the conversation; a separate successful backup result was not captured. Verify that your own backup exists before modifying the file. The inspection was captured and located this guard at lines 224–225 in my installed build.
VNA configuration file — inspected Python excerpt, not a command to execute:
if "AMD" in vendor_info and "AMD EPYC" not in model_name:
self.error_exit("Unsupported CPU: %s" % model_name)VNA root shell — exact recorded edit, verification, and reboot:
sed -i '224,225s/^/# /' /opt/vmware/nsx-edge/bin/config.py
sed -n '224,225p' /opt/vmware/nsx-edge/bin/config.py
rebootDo not apply those line numbers blindly to another version. Find the guard in that installed file and confirm that only the intended two lines will be commented. The screenshot shows both lines prefixed with #. The first sed command changes the file; the second prints the edited lines so you can verify them.
Wait for the VNA deployment to recover #
After applying the workaround and rebooting, I returned to Configure → Networking → VNA Clusters. The node briefly showed L2 Config Failed, then In Progress. I left the deployment running; it reached Up / Success after roughly five minutes.
With the cluster and node Up / Success, I continued. BFD Sessions: Not Available remained visible. This lab used a single VNA node.
7. Create the two subscribed content libraries #
The Supervisor image source and the guest Kubernetes release source serve different consumers. I used separate names and subscriptions:
| Library | Subscription URL | Purpose |
|---|---|---|
Supervisor-Subscribed-Library |
https://wp-content.broadcom.com/supervisor/v1/latest/lib.json |
Supervisor lifecycle images |
VKR-Subscribed-Library |
https://wp-content.broadcom.com/v2/latest/lib.json |
VMware Kubernetes Releases for guest-cluster provisioning |
In nested vCenter → Content Libraries → Create, name the library, select the nested vCenter, choose Subscribed content library, enter its URL, and select vsanDatastore. Use automatic synchronization, no authentication for these public subscriptions, and download content immediately. Wait for the needed images to finish downloading before using them.
Downloading immediately uses storage and transfer time up front so the images are available locally.
The local Fleet-gateway subscription I tried could not fetch its items in this build:
Content Library subscription field — local URL that failed in my lab:
https://vcf-flt01.vcf.lab.devynharrington.com/depot-service/content-gateway/PROD/COMP/VKR/lib.jsonI used the public VKR subscription instead, which worked in my lab.
Broadcom’s library-association guidance places the Supervisor image-source association under Supervisor Management → Content Distribution → Supervisor Images Library. The 9.1 Supervisor library article confirms the public Supervisor endpoint.
After Supervisor is running, associate VKR-Subscribed-Library under Supervisor → Configure → General → Kubernetes Service, then make it available through the namespace’s VM Service / Content Libraries configuration, as shown in section 10.
8. Activate a Medium Supervisor #
With the VNA healthy and image sources prepared, I opened Supervisor Management → Add Supervisor, chose VCF Networking with VPC, and selected the zone containing vcf-mgmt-cl01.
| Supervisor setting | Recorded value |
|---|---|
| Name | vcf-mgmt-supervisor01 |
| Cluster / vCenter | vcf-mgmt-cl01 / vc01.vcf.lab.devynharrington.com |
| Control-plane VM count | One |
| Size | Medium; UI showed 8 vCPU, 24 GB memory, 48 GB storage |
| Management assignment | DHCP |
| Management network | vcf-mgmt-cl01-vds01-pg-vm-mgmt |
| Management floating IP | 192.168.88.91 |
| Management DNS / search domain | 192.168.88.1 / vcf.lab.devynharrington.com |
| Management NTP | time.google.com |
| NSX project | Default |
| VPC connectivity profile | Default VPC Connectivity Profile |
| External block | VCF-External-IP-Block; 10.1.35.0/24 |
| Workload DNS / NTP | 192.168.88.1 / time.google.com |
| API Server DNS Name | vcf-mgmt-supervisor01.vcf.lab.devynharrington.com |
| Control-plane, ephemeral, and image storage policies | vcf-mgmt-cl01 - Optimal Datastore Default Policy - AutoRAID |
DHCP for the control-plane management configuration and the explicit .91 floating address are separate fields. Ensure DHCP and routing are available on the selected management network; copying the floating IP alone does not supply them.
Initial Supervisor sizing before my Medium selection
I chose Medium for the planned Kubernetes use. The UI warned that it could not later scale down. Keep these sizes separate: Simple/Small fleet, Large VNA, Medium Supervisor, and later best-effort-small VKS node VMs. They size different components.
Review the network, storage, sizing, and endpoint choices and activate the Supervisor. Entering the API DNS name does not itself prove that a DNS A record was created. The resulting API address was 10.1.35.7; its current DNS answer remains an uncaptured check.
9. Follow Supervisor reconciliation through to both Running states #
The guest, management-network, and Kubernetes control-plane conditions completed while workload networking and service configuration were still in progress.
During reconciliation, I observed the following errors and status changes:
| Observation | What I could conclude |
|---|---|
| NSX trust-management/token-principal-identities request returned HTTP 503 | An NSX/API dependency was failing at that moment. |
| Pinniped HTTPS proxy synchronization and container discovery failed | The authentication-related workload-network configuration was incomplete. |
| Outer inventory reported zero free GHz | The physical machine was under CPU pressure; this was not a CPU Ready measurement. |
| NSX showed Stable / Up, repository sync, roughly 96% CPU, four alarms | NSX remained running under high CPU load. |
| Supervisor Config Running, Host Config Configuring | Control-plane progress had preceded host configuration. |
Apply Solution health check failed for nested-esx03 |
A host health check blocked that reconciliation attempt. |
NSX, Pinniped, and CPU-pressure evidence during reconciliation
Delayed host conditions and Apply Solution health checks
For a similar delay, inspect Supervisor Management → status details, the cluster’s recent tasks, and cluster → Monitor → vSAN → Skyline Health. Resolve failures; silence only understood warnings related to nested hardware.
I did not isolate a single cause of the recovery.
Both Config Status and Host Config Status reached Running, while the Supervisor remained Medium:
The UI showed three hosts, three system namespaces, API address 10.1.35.7, and Supervisor version v1.34.9+vmware.1-vsc.9.1.1.0-25712839. Both Running columns—not just the existence of a control-plane VM—were the checkpoint for continuing.
Mac terminal — suggested DNS verification, not a captured completed test:
nslookup vcf-mgmt-supervisor01.vcf.lab.devynharrington.com 192.168.88.1The expected target for this lab is 10.1.35.7. My later CLI used the API IP directly.
10. Configure the namespace and its Public SubnetSet #
At Supervisor → Configure → General → Kubernetes Service, I associated VKR-Subscribed-Library.
In Supervisor Management / Namespaces → Create Namespace, I selected vcf-mgmt-supervisor01, named the namespace homelab-vks-ns, and selected the zone containing the cluster. I left Override default network settings unchecked, then reviewed and created the namespace.
Then I configured the namespace resources:
| Namespace item | Setting or check |
|---|---|
| Permissions | Grant the intended principal Can Edit. The final permission dialog is not captured. |
| VM class | best-effort-small: 2 vCPU, 4 GiB, no CPU/memory reservation shown. |
| Storage | vcf-mgmt-cl01 - Optimal Datastore Default Policy - AutoRAID. |
| VM Service / Content Libraries | Associate VKR-Subscribed-Library here as well as under the Supervisor’s Kubernetes Service. |
| Resources | Confirm the Kubernetes and Network service pages load. |
If the VKS wizard cannot find the VM class, release, or storage, check these namespace associations before recreating the library or cluster.
Create public before creating VKS
#
In homelab-vks-ns → Resources → Network service → SubnetSets → New SubnetSet, I used these settings:
| Field | Final setting |
|---|---|
| Name | public |
| Access Mode | Public |
| Auto-allocate IP CIDR | On |
| Size | 32 IPs |
| DHCP | None |
| Labels | None shown |
The 32 IPs setting requests a /27; fewer addresses are available for guest nodes. The screenshot does not show the allocated CIDR.
Public is a VPC access mode, not public internet exposure; these remain private lab addresses. The default private network was rejected without SNAT (source network address translation), so I used public.
11. Create VKS with Custom Configuration and verify it from Supervisor #
In Resources → Kubernetes service → Create, I selected Custom Configuration to choose public as the primary node network.
General settings and storage #
| Field | Setting |
|---|---|
| Cluster name | kubernetes-cluster-dhby |
| Type / configuration | Cluster API / Custom Configuration |
| ClusterClass | builtin-generic-v3.7.0 |
| Kubernetes release | v1.36.2---vmware.2-vkr.3 |
| VM class | best-effort-small |
| Node VM storage class | vcf-mgmt-cl01-optimal-datastore-default-policy-autoraid |
| Default PVC storage class | vcf-mgmt-cl01-optimal-datastore-default-policy-autoraid-latebinding |
| Additional node volumes | None shown |
| Custom storage annotations / labels | Blank |
Cluster API manages the cluster using a ClusterClass topology. A PVC (PersistentVolumeClaim) requests application storage; its default storage class is separate from the node VM storage class.
I kept the generated name kubernetes-cluster-dhby in all commands. homelab-vks01 refers to the earlier failed deployment.
Primary network and replicas #
| Field | Final value |
|---|---|
| CNI / pod CIDR | Antrea / 10.95.0.0/16 |
| Primary interface | eth0 |
| Network selection | Use custom primary network |
| Network | public (Subnet Set) |
| Primary-interface MTU | 1500 |
| Additional static routes / interfaces | None |
| Control plane | One replica; Photon 5; Overrides off |
| Worker pool | One replica; Photon 5; advanced settings unchecked |
| Generated pool name | kubernetes-cluster-dhby-np-vsw3 |
Intermediate primary-network form before the MTU correction
If you do not change the network to public and leave the default private network selected, you will get this error at step 5, Review and Confirm, in this VPC namespace without SNAT:
I set the guest primary-interface MTU to 1500 because jumbo frames were unverified across the external VLAN path. The separate VDS, vMotion, and vSAN settings remained 9000.
The default YAML showed Service CIDR 10.96.0.0/12 and domain cluster.local; I did not verify these against the full runtime configuration.
With public selected, I reviewed and created the cluster, then confirmed its status was Available:
Authenticate to Supervisor and select its namespace #
The VMware Supervisor CLI reference distinguishes three context scopes:
| Context | Resources |
|---|---|
| Supervisor | Supervisor infrastructure and namespace resources |
Supervisor namespace supervisor-redeploy:homelab-vks-ns |
Guest Cluster API object, node VMs, and namespace network resources |
| Guest cluster | Kubernetes nodes, system pods, and application namespaces |
Mac terminal — create the Supervisor context:
vcf context create supervisor-redeploy \
--endpoint https://10.1.35.7 \
--username [email protected] \
--auth-type basic \
--type k8s \
--insecure-skip-tls-verifyEnter the SSO password at the prompt. This lab command skips TLS certificate verification; omit --insecure-skip-tls-verify when your client trusts the endpoint certificate.
Login succeeded and the contexts were saved despite a Darwin ARM64 telemetry:v9.0.2 lookup returning MANIFEST_UNKNOWN.
Mac terminal — Supervisor namespace context:
vcf context use supervisor-redeploy:homelab-vks-ns
kubectl config current-context
kubectl get clusters -n homelab-vks-ns| Cluster result | Observed output |
|---|---|
| Name / ClusterClass | kubernetes-cluster-dhby / builtin-generic-v3.7.0 |
| Available / Phase | True / Provisioned |
| Control plane desired / available / up-to-date | 1 / 1 / 1 |
| Worker desired / available / up-to-date | 1 / 1 / 1 |
| Version | v1.36.2+vmware.2 |
A Harbor plugin-source discovery warning also appeared, but the cluster query succeeded. The warning remained unresolved; I did not need to install Harbor for this query.
Both guest-node VMs were powered on, though the nested ESXi hosts still showed warnings.
If the cluster is still reconciling, use these optional diagnostics; not all were run during this deployment.
Mac terminal — optional diagnostics in the Supervisor namespace context:
kubectl get clusters -n homelab-vks-ns -w
kubectl get machines -n homelab-vks-ns -o wide
kubectl get virtualmachines -n homelab-vks-ns -o wide
kubectl get events -n homelab-vks-ns --sort-by=.lastTimestamp
kubectl describe subnetset public -n homelab-vks-ns
kubectl get subnets -n homelab-vks-ns -o wide
kubectl get cluster kubernetes-cluster-dhby -n homelab-vks-ns -o yamlStop the watch with Ctrl+C, then run the remaining commands. Cluster creation continues.
12. Enter the guest cluster with its downloaded kubeconfig #
With VKS Available, I connected to the guest Kubernetes API to verify its nodes and run a workload. The Supervisor context manages the VKS cluster object; the guest context accesses its nodes, pods, and Services. The vSphere namespace homelab-vks-ns and guest namespace lab-validation belong to different clusters.
From supervisor-redeploy:homelab-vks-ns, I tried switching to the expected guest context:
Mac terminal — switch to the guest context:
vcf context use supervisor-redeploy:homelab-vks-ns:kubernetes-cluster-dhby
kubectl config current-contextThe guest context was not found. A telemetry:v9.0.2 lookup also returned MANIFEST_UNKNOWN, but that does not establish why the context was missing. I checked the available contexts:
Mac terminal — inspect VCF contexts:
vcf context listInstead, I downloaded the kubeconfig from nested vCenter → homelab-vks-ns → Resources → Kubernetes Service → kubernetes-cluster-dhby → Download Kubeconfig File. Keep this credential file private and use your own cluster name and file path.
Mac terminal — locate the downloaded file:
ls -lt ~/Downloads/*kube*I first used --kubeconfig to query the guest cluster explicitly:
Mac terminal — guest-cluster access with explicit kubeconfig:
kubectl --kubeconfig="$HOME/Downloads/kubernetes-cluster-dhby-kubeconfig.yaml" config get-contexts
kubectl --kubeconfig="$HOME/Downloads/kubernetes-cluster-dhby-kubeconfig.yaml" get nodes -o wide
kubectl --kubeconfig="$HOME/Downloads/kubernetes-cluster-dhby-kubeconfig.yaml" get pods -A -o wideThe context was kubernetes-cluster-dhby-admin@kubernetes-cluster-dhby. After the queries succeeded, I selected this kubeconfig for the session:
Mac terminal — select the guest cluster for subsequent commands:
export KUBECONFIG="$HOME/Downloads/kubernetes-cluster-dhby-kubeconfig.yaml"All kubectl commands below use this export; repeat it in a new terminal. It leaves the default kubeconfig unchanged and bypasses the missing VCF guest context.
13. Check nodes, system pods, storage classes, and resource use #
I checked guest-cluster readiness before deploying the test workload.
Mac terminal — guest-cluster context:
kubectl get nodes -o wide
kubectl get pods -A -o wide
kubectl get storageclass
kubectl top nodes| Observation | Control plane | Worker |
|---|---|---|
| Status | Ready | Ready |
| Internal IP | 10.1.35.66 |
10.1.35.67 |
| Kubernetes | v1.36.2+vmware.2 |
v1.36.2+vmware.2 |
| CPU at capture | 499m / 25% | 76m / 3% |
| Memory at capture | 2020Mi / 71% | 510Mi / 18% |
Both nodes ran Photon OS on amd64. The worker’s role column showed <none>; its node-pool configuration identified it as the worker. Generated node names change when the cluster is recreated.
The system-pod query showed 20 pods Running, with all listed containers ready, summarized below. Figure A05 shows the other checks.
| System component | Pods and ready containers | Placement | Observed restarts |
|---|---|---|---|
| Antrea agents | 2 pods, each 2/2 | One per node | 0 / 0 |
| Antrea controller | 1/1 | Control plane | 0 |
| CoreDNS | 2 pods, each 1/1 | Control plane | 0 / 0 |
| etcd | 1/1 | Control plane | 0 |
| Kubernetes API server | 1/1 | Control plane | 0 |
| Kubernetes controller manager | 1/1 | Control plane | 1 |
| kube-proxy | 2 pods, each 1/1 | One per node | 0 / 0 |
| Kubernetes scheduler | 1/1 | Control plane | 1 |
| metrics-server | 1/1 | Control plane | 0 |
| snapshot-controller | 1/1 | Control plane | 0 |
| secretgen-controller | 1/1 | Control plane | 0 |
| kapp-controller | 2/2 | Control plane | 0 |
| guest-cluster-auth-svc | 1/1 | Control plane | 0 |
| guest-cluster-cloud-provider | 1/1 | Control plane | 1 |
| vSphere CSI controller | 7/7 | Control plane | 0 |
| vSphere CSI nodes | 2 pods, each 3/3 | One per node | 4 control plane / 1 worker |
Some pods had earlier restarts, but none showed an active crash loop. Antrea provides networking, CoreDNS handles DNS, metrics-server reports resource use, and vSphere CSI integrates storage. The workload below tests these functions beyond pod readiness.
I had two AutoRAID storage classes:
| Storage class | Binding | Default |
|---|---|---|
vcf-mgmt-cl01-optimal-datastore-default-policy-autoraid |
Immediate |
No |
vcf-mgmt-cl01-optimal-datastore-default-policy-autoraid-latebinding |
WaitForFirstConsumer |
Yes |
Both use csi.vsphere.vmware.com, support volume expansion, and reclaim volumes with Delete. Immediate binds or provisions storage when the PVC is created; WaitForFirstConsumer waits for a pod’s scheduling requirements. The manifest below uses the default class, so check yours before applying it. See the storage-class documentation.
Checkpoint: nodes and system containers are ready, storage is available, and metrics work. Resource figures are snapshots, not capacity benchmarks.
14. Create the BusyBox and PVC validation workload #
I used a BusyBox pod and PVC in the guest namespace lab-validation to test image retrieval, networking, and storage.
Paste the entire block into the guest-cluster terminal, including the closing EOF. Scroll to view the manifest; Copy copies the entire block.
Mac terminal — guest-cluster context:
kubectl apply -f - <<'EOF'
apiVersion: v1
kind: Namespace
metadata:
name: lab-validation
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: test-storage
namespace: lab-validation
spec:
accessModes:
- ReadWriteOnce
resources:
requests:
storage: 1Gi
---
apiVersion: v1
kind: Pod
metadata:
name: lab-test
namespace: lab-validation
spec:
securityContext:
runAsNonRoot: true
runAsUser: 1000
runAsGroup: 1000
fsGroup: 1000
seccompProfile:
type: RuntimeDefault
containers:
- name: test
image: busybox:1.37
command: ["sh", "-c", "sleep 86400"]
securityContext:
allowPrivilegeEscalation: false
capabilities:
drop: ["ALL"]
resources:
requests:
cpu: 10m
memory: 16Mi
limits:
cpu: 100m
memory: 64Mi
volumeMounts:
- name: data
mountPath: /data
volumes:
- name: data
persistentVolumeClaim:
claimName: test-storage
EOFThe pod mounts a 1 GiB, ReadWriteOnce volume at /data, using the default latebinding AutoRAID class. It runs as nonroot UID/GID 1000 with fsGroup: 1000, RuntimeDefault seccomp, no privilege escalation, and all Linux capabilities dropped. See the security-context documentation.
I waited for readiness, then checked the pod and PVC:
Mac terminal — guest-cluster context:
kubectl wait -n lab-validation --for=condition=Ready pod/lab-test --timeout=180s
kubectl get pods,pvc -n lab-validation -o wideThe wait timed out with lab-test in ImagePullBackOff, while test-storage was Bound. I inspected the pod to isolate the image-pull failure:
Mac terminal — inspect the failing pod:
kubectl describe pod lab-test -n lab-validationEvents showed SuccessfulAttachVolume: scheduling and volume attachment had succeeded. The image had not been pulled, so I investigated DNS before testing writes to the volume.
15. Diagnose the image-pull timeout and permit lab DNS on MikroTik #
In the Events section at the bottom of the kubectl describe pod lab-test -n lab-validation output from the previous step, I found this Docker Hub DNS timeout:
Failed to pull image "busybox:1.37":
failed to pull and unpack image "docker.io/library/busybox:1.37":
failed to resolve image: failed to do request:
Head "https://registry-1.docker.io/v2/library/busybox/manifests/1.37":
dial tcp: lookup registry-1.docker.io on 127.0.0.53:53:
read udp 127.0.0.1:44928->127.0.0.53:53: i/o timeoutThe node’s local resolver stub, 127.0.0.53, timed out resolving Docker Hub before connecting to the registry. The error does not identify the upstream resolver or indicate an authentication or rate-limit failure.
I inspected the MikroTik configuration. Use your router’s address and login:
Mac terminal — connect to MikroTik:
MikroTik RouterOS terminal — inspect DNS, firewall, interfaces, and NAT:
/ip/dns/print
/ip/firewall/filter/print detail
/interface/list/member/print
/ip/firewall/nat/print detailDNS was already enabled with allow-remote-requests=yes and dynamic upstream servers 206.225.75.226 and 206.225.75.225. WAN masquerade was also configured; its presence alone did not verify internet access.
The input rule defconf: drop all not coming from LAN matched in-interface-list=!LAN, and vlan2006-external was missing from the LAN list:
Queries to the router use its input chain; forwarding and NAT rules do not override it. I allowed DNS from 10.1.35.0/24 on vlan2006-external before the input drop.
Run each line once, after checking the interface, subnet, and drop-rule comment against your router. The place-before selector must match the intended drop rule. DNS needs both UDP and TCP port 53; see MikroTik’s DNS documentation.
MikroTik RouterOS terminal — allow lab DNS:
/ip/firewall/filter/add chain=input action=accept in-interface=vlan2006-external src-address=10.1.35.0/24 protocol=udp dst-port=53 place-before=[find where comment="defconf: drop all not coming from LAN"] comment="VKS lab DNS UDP"
/ip/firewall/filter/add chain=input action=accept in-interface=vlan2006-external src-address=10.1.35.0/24 protocol=tcp dst-port=53 place-before=[find where comment="defconf: drop all not coming from LAN"] comment="VKS lab DNS TCP"With DNS allowed only from the lab interface and subnet, I watched the existing pod retry its image pull.
Mac terminal — guest-cluster context:
kubectl get pod lab-test -n lab-validation -wThe pod reached 1/1 Running with zero restarts. Press Ctrl+C to stop the watch. If your pod remains in backoff, run kubectl describe pod lab-test -n lab-validation again and check the latest error.
Recovery after the firewall change supports the DNS diagnosis. The BusyBox pull and later NGINX rollout were functional checks; I did not capture firewall counters or a packet trace.
16. Prove cluster DNS and a write/read on the mounted volume #
With BusyBox running, I tested cluster DNS:
Mac terminal — run inside the guest test pod:
kubectl exec -n lab-validation lab-test -- nslookup kubernetes.default.svc.cluster.localDNS server 10.96.0.10 resolved kubernetes.default.svc.cluster.local to 10.96.0.1. BusyBox ran on the worker and CoreDNS on the control plane, exercising DNS across nodes. This tested an internal name, not an external-domain lookup.
These guest addresses do not verify the full Service CIDR or resolve the earlier Supervisor CIDR discrepancy.
Next I wrote a file to /data and read it back.
Mac terminal — run inside the guest test pod:
kubectl exec -n lab-validation lab-test -- sh -c 'echo "VCF 9.1.1 VKS storage test passed" > /data/test.txt && cat /data/test.txt'The bound PVC, successful attachment, and write/read confirmed provisioning and basic storage I/O. Persistence across pod replacement, failure recovery, and performance were not tested.
17. Deploy NGINX and open it through a LoadBalancer #
I deployed one NGINX replica in lab-validation with a LoadBalancer Service for access from my Mac. It does not use the BusyBox PVC.
The unprivileged NGINX image listens on 8080 and supports the nonroot settings below. I used the moving tag stable-alpine; its digest was not captured.
Paste the entire block, including EOF, into the guest-cluster terminal. The lab-validation namespace must already exist. Scroll to view the full manifest; Copy copies the entire block.
Mac terminal — guest-cluster context:
kubectl apply -f - <<'EOF'
apiVersion: apps/v1
kind: Deployment
metadata:
name: lab-web
namespace: lab-validation
spec:
replicas: 1
selector:
matchLabels:
app: lab-web
template:
metadata:
labels:
app: lab-web
spec:
securityContext:
runAsNonRoot: true
runAsUser: 101
seccompProfile:
type: RuntimeDefault
containers:
- name: nginx
image: nginxinc/nginx-unprivileged:stable-alpine
ports:
- containerPort: 8080
securityContext:
allowPrivilegeEscalation: false
capabilities:
drop: ["ALL"]
resources:
requests:
cpu: 25m
memory: 32Mi
limits:
cpu: 250m
memory: 128Mi
readinessProbe:
httpGet:
path: /
port: 8080
initialDelaySeconds: 5
periodSeconds: 5
---
apiVersion: v1
kind: Service
metadata:
name: lab-web
namespace: lab-validation
spec:
type: LoadBalancer
selector:
app: lab-web
ports:
- name: http
port: 80
targetPort: 8080
EOFThe Service selects app: lab-web and forwards 80/TCP → 8080. The pod runs as nonroot UID 101 with restricted privileges and an HTTP readiness probe.
I checked the rollout:
Mac terminal — guest-cluster context:
kubectl rollout status deployment/lab-web -n lab-validation --timeout=180sAfter the rollout succeeded, I checked the Service address:
Mac terminal — guest-cluster context:
kubectl get service lab-web -n lab-validation -w| Service field | Recorded result |
|---|---|
| Name / namespace | lab-web / lab-validation |
| Type | LoadBalancer |
| ClusterIP | 10.99.227.208 |
| External IP | 10.1.35.34 |
| Service output | 80:30703/TCP |
| Browser-facing Service port | 80 |
| Container target port | 8080 |
| Allocated NodePort | 30703 |
Press Ctrl+C once the address appears, then open your Service’s assigned IP. Check it again if you recreate the Service.
Mac browser — open the assigned address:
http://10.1.35.34The application path was Mac browser → 10.1.35.34:80 → lab-web Service → NGINX pod on 8080. Browser access uses Service port 80, not NodePort 30703.
NGINX was reachable from my home network through the VKS LoadBalancer. This tested private lab HTTP access; public hosting and HTTPS were outside the scope.
18. Results, practical lessons, and the next application #
The rebuild reached these milestones:
| Milestone | Confirmed result |
|---|---|
| VCF 9.1.1 | Installer deployment completed with Day-0 Automation disabled |
| VNA | Large vna01 reached Up / Success after the Ryzen CPU-guard edit and recovery |
| Supervisor | Medium Supervisor and all host configurations Running |
| Namespace and network | homelab-vks-ns with Public SubnetSet public, 32 IPs, and /24 external networking |
| VKS | Cluster Provisioned and Available; downloaded kubeconfig provided access |
| Guest readiness | Both nodes Ready; 20 system pods Running with containers ready |
| DNS and image retrieval | Internal DNS worked; BusyBox and NGINX images pulled after the router DNS fix |
| Storage | 1 GiB PVC bound and attached; /data/test.txt write/read passed |
| Application | NGINX served through 10.1.35.34:80 |
This was a single-physical-host nested lab with one Supervisor control-plane VM and one VKS control plane and worker. It demonstrated basic functionality, not production availability, performance, or recovery.
Future work: a custom application #
Next, I plan to deploy a custom application on this VKS cluster using an AMD64-compatible container image.
19. References #
Sources for the deployment workflow and supporting concepts: