Featured image of post Asymmetric routing after enabling VPC

Asymmetric routing after enabling VPC

The VPC and Supervisor were running, but access to the Supervisor API server was unpredictable. I could reach the portal only 6 out of 10 times. I needed a stable connection before I could configure the endpoint with VCF-CLI.

After checking IP configuration, DNS, and TLS, I narrowed the problem down to routing.

The problem:

I observed that access to the Supervisor API (10.31.x.x:443) was flapping:

  • sometimes it worked

  • other times it hung, and I could see SYN retransmissions with no response

I started a packet capture on the pfSense interface for my management subnet. I filtered on the Supervisor API IP, TCP, and port 443, then ran curl repeatedly from PowerShell to generate the traffic.

PS C:\Users\Vidar> 1..30 | % { curl.exe –ssl-no-revoke -I https://10.x.x.x/api | Out-Null; Start-Sleep -Milliseconds 400 }

Wireshark showed exactly what was happening:

  • the client on mgmt subnet(10.0.x.x) sent a SYN

  • but did not receive a SYN/ACK back after several attempts

  • then suddenly it would work again on a new connection

Why it happened:

On the NSX Tier-0 (T0), there was no specific route for the management network:

Mgmt CIDR → pfSense

Because of that, when T0 needed to send return traffic back to my client on the mgmt subnet, it had to rely on the default route.

The issue was that T0 had two default routes at the same time, with equal preference:

  • 0.0.0.0/0 → uplink toward pfSense

  • 0.0.0.0/0 → 169.254.2.3 (automatically injected from the Transit Gateway / VPC)

When two default routes have the same preference, NSX will often use ECMP or per-flow hashing:

  • some flows returned through 10.0.x.x(uplink toward pfSense) and worked correctly

  • some flows returned through 169.254.2.3, which was the wrong path for management traffic

  • as a result, the SYN/ACK never made it back to the client

That is why the behavior appeared random.

What fixed it:

The fix was to add a specific route on T0:

Mgmt CIDR → Uplink to pfSense

I added this in the same place as the default route on my T0: a static route for the management subnet CIDR with pfSense as the next hop.

This forces a longest-prefix match:

  • traffic destined for mgmt subnet will always use the /24 route

  • it will never use either of the default routes

Once that route was added, the flapping disappeared.

Where did the “unknown” default route come from?

I did not manually configure 0.0.0.0/0 → 169.254.2.3.

It was automatically learned from the Transit Gateway (TGW) associated with the NSX VPC/Project setup.

That route is intended for VPC/TGW functionality, but it indirectly affected this setup because there was no specific route for the management network.

What dynamic routing (BGP/OSPF) would have done:

If BGP or OSPF had been configured between pfSense and T0, pfSense could have advertised internal networks such as the management CIDR.

In that case, T0 would automatically have learned a specific route for the management network, and no manual static route would have been needed.

The key point is that OSPF or BGP are not the solution by themselves. What actually fixes the issue is giving the Tier-0 a specific route for the internal network, so return traffic does not rely on the default route.

Before the VPC was introduced, the Tier-0 had a single default route toward the firewall, so return traffic followed one consistent path and the setup worked. The issue only appeared after the VPC introduced a competing default route, which made the return path ambiguous and caused asymmetric routing. Adding a specific route for the internal subnet — either statically or learned dynamically through BGP/OSPF — restored a stable return path.