An expired NSX Manager certificate is not a reason to start in the UI. The certificate the browser is warning on is the one you are there to replace, and the API calls that replace it cannot verify the thing they are about to throw away. The job is a plan, then a write: one certificate per manager node, the cluster VIP last, the data plane left alone.
The Challenge
After install, the manager nodes and the cluster VIP have self-signed certificates. The API certificate and the VIP certificate are valid for 825 days on a new deployment. They do not refresh themselves on an upgrade. An estate that has been left alone hits that date, and the first person who opens the manager gets a warning instead of a procedure.
A single certificate that covers every node and the VIP is one supported path. It is not the path this tool takes. This tool gives each node its own certificate, with that node’s name and address as the subject alternative names, and a separate certificate for the VIP. NSX will not accept a certificate that is also a CA and has a key. The tool sets that constraint when it generates the certificate, because a CA certificate fails the import for a reason that does not mention the key.
What the Tool Does
Without --commit it reads. It lists the manager nodes, prints the certificate each one is serving, and says whether that certificate is already expired. It writes nothing.
With --commit it makes a self-signed certificate, 825 days, imports it, and applies it. Node certificates use the API service type and that node’s id. The VIP certificate uses the management-cluster service type and is applied last. After each apply it waits until that manager answers again. The API on a node is unavailable for about a minute. The data plane is not touched.
If a self-signed certificate is not acceptable, --csr-dir writes one request and key per name and stops. --import-dir applies the signed certificates from there. Failed edge deployments are a separate flag. Name them, or they are not deleted. A delete that the API refuses is retried with force. That is not part of the certificate replacement. It is in the same tool because the same quiet estate often has both.
The Results
The plan prints the names, the addresses, and the dates before anyone generates a key. After a commit, each manager and the VIP serve a certificate whose name matches that node, with a new end date. Keys land in a directory that is not world-readable. Nothing secret is printed.
Lessons Learned
Verifying the certificate you are there to replace will fail, because that certificate is already expired. The tool turns verification off for these calls and nowhere else. Applying every node at once, or the VIP before the nodes, is how you lose the cluster you are using to do the work. One node, then the next, then the VIP.
VCF 9.1.1 also has a shortcut in the manager UI for certificates that are expired or inside 30 days. Use it when the UI is the thing you trust. Use the API when you want a plan you can read before the first key is generated, or when the browser warning is the reason you are not clicking through the UI.
Getting Started
Run it beside this. The names are the essential.coach lab. Replace them with yours. The first command writes nothing.
export NSX_PASSWORD='...'
python3 nsx_cert_fix.py --manager nsx.essential.coach
python3 nsx_cert_fix.py --manager nsx.essential.coach --commit
python3 nsx_cert_fix.py --manager nsx.essential.coach --delete-edges edge01,edge02 --commit
Clone nsx-cert-fix. Export NSX_PASSWORD, or let it prompt. Run it against the manager VIP with no --commit and read the table. Commit only the managers the table marked expired. If the site requires a CA, stop at --csr-dir and do not generate the self-signed set.
Conclusion
An expired manager certificate is a replacement, not a reboot and not a password change. Read the dates. Replace the node certificates one at a time. Replace the VIP last. Leave the data plane alone.
References
Replace NSX Certificates Using API is the procedure this tool follows. After install, the manager nodes and the cluster have self-signed certificates. A node certificate is applied with service_type=API and that node’s id. The VIP certificate is applied with service_type=MGMT_CLUSTER. The same page covers Local Manager and Global Manager API certificates when federation is in use. The 825-day default for those API and VIP certificates is in Certificates for NSX and NSX Federation. The UI path, including the 9.1.1 shortcut for certificates that are expired or inside 30 days, is Replace NSX Certificates from NSX Manager.
Repository: github.com/noahfarshad/nsx-cert-fix
Related Stories:
- Moving NSX DFW Between Global Managers Without Turning the Rules On — another NSX API job that stays outside the VCF buttons
- NSX Federation Made the Same Segment Appear Twice — Global Manager and Local Manager, when the certificate you are replacing is one of those
