Enforcing mTLS Between Microservices: Stop Internal MITM
Lock down east–west traffic with mutual TLS. Learn production mTLS setup, three common handshake failures Australian teams hit, and DIY fixes—plus when to book Fixwebnode for a remote zero-trust pipeline.
If a hacker breaches a single application node on your network, they should not be able to snoop on internal traffic. This guide walks Australian sole traders, small businesses, and ops teams through enforcing mutual TLS (mTLS) between microservices so every hop authenticates both sides and encrypts the payload end to end.
Plain HTTP or one-way TLS inside the VPC is still a man-in-the-middle risk: a compromised pod, misrouted container, or rogue process on the same segment can read tokens, PII, and service credentials. Fixwebnode works remotely with teams across Australia to design and harden that encrypted zero-trust pipeline—certificate authorities, workload identity, and verification—without turning your stack into a freelance bidding board. Start with the steps below, or open a conversation via Fixwebnode website repair Australia when you need a specialist pair of eyes on production.
Why mTLS matters for the man-in-the-middle network fix
Mutual TLS means the client proves its identity with a certificate and the server does the same. Unlike edge HTTPS alone, mTLS protects service-to-service calls: API gateways to auth, workers to queues, and multi-tenant backends that share a cluster. For production in Australia—whether you run Kubernetes on a cloud account or bare VMs—you need a private CA (or mesh CA), short-lived leaf certs, strict verification, and a rollout plan that never leaves half the fleet on cleartext.
Remote delivery fits this work: diagnostics, cert inventory, config reviews, and cutovers can be done over secure sessions. Coverage and how we engage by region is summarised on our service areas page.
Why does mTLS fail between microservices in Australian production?
Most failures are not “TLS is broken”—they are trust-store gaps, hostname or SPIFFE identity mismatches, or clock skew that makes valid certs look expired. Fix the trust chain and identity first; only then chase cipher suites or mesh sidecars.
| Symptom | Quick fix | When to call Fixwebnode |
|---|---|---|
| handshake: unknown CA / unable to get local issuer | Install the issuing CA (not only the leaf) on every peer | Multiple CAs, rotation mid-flight, or mixed mesh and bare TLS |
| certificate verify failed: name mismatch | Align SAN/CN or URI SAN with the dialled name | SPIFFE IDs, multi-tenant hostnames, or custom resolvers |
| works in staging, intermittent 5xx in prod | Check NTP skew and cert NotBefore/NotAfter windows | Fleet-wide time drift or automated renewals colliding |
Common mTLS issues (unique symptoms)
1. Incomplete CA chain on one side of the call
Clients connect, then abort with x509: certificate signed by unknown authority or OpenSSL unable to get local issuer certificate. The leaf cert is present, but the intermediate or root that signed it is missing from the peer’s trust store—classic after a CA rotation or when only the server cert was copied into a secret.
2. SAN / identity mismatch (hostname or SPIFFE URI)
Logs show certificate is valid for …, not … or Envoy/Istio TLS error: 268435581. The cert is trusted, but the name the client dials (service DNS, pod FQDN, or spiffe:// ID) is not listed in Subject Alternative Name. Partial mesh enablement makes this worse: some pods use SPIFFE, others still use DNS SANs.
3. Clock skew and “not yet valid” / expired leaf certs
Handshakes fail only on certain nodes or after deploys. openssl verify on the box that fails reports the cert is not yet valid or has expired while a neighbouring node succeeds. Short-lived mTLS certs (hours to a few days) amplify even a few minutes of NTP drift.
4. Client certificate never sent (one-way TLS by accident)
Server access logs show successful TLS but no peer certificate; authorisation still relies on network location. Root cause is usually ssl_verify_client off, missing ssl_client_certificate, or a client library that only loads a CA file and not a key pair.
How to run mutual TLS in production (baseline setup)
Prerequisites: Linux hosts or Kubernetes nodes with OpenSSL 3.x (or 1.1+), root or workload identity to mount secrets, and a private CA you control (step-ca, CFSSL, HashiCorp Vault PKI, or your mesh CA). Do not reuse public website certs for internal mTLS.
Step 1 — Create a dedicated internal CA and leaf pair (lab or bootstrap)
mkdir -p ~/mtls-lab/{ca,server,client} && cd ~/mtls-lab
openssl genrsa -out ca/ca.key 4096
openssl req -x509 -new -nodes -key ca/ca.key -sha256 -days 3650 \
-subj "/CN=Internal Services CA/O=YourOrg/C=AU" -out ca/ca.crt
openssl genrsa -out server/server.key 2048
openssl req -new -key server/server.key -subj "/CN=orders.internal" \
-addext "subjectAltName=DNS:orders.internal,DNS:orders.default.svc,DNS:orders.default.svc.cluster.local" \
-out server/server.csr
openssl x509 -req -in server/server.csr -CA ca/ca.crt -CAkey ca/ca.key -CAcreateserial \
-out server/server.crt -days 90 -sha256 -copy_extensions copy
openssl genrsa -out client/client.key 2048
openssl req -new -key client/client.key -subj "/CN=checkout-client" \
-addext "subjectAltName=DNS:checkout.internal,URI:spiffe://cluster.local/ns/default/sa/checkout" \
-out client/client.csr
openssl x509 -req -in client/client.csr -CA ca/ca.crt -CAkey ca/ca.key -CAcreateserial \
-out client/client.crt -days 90 -sha256 -copy_extensions copyProduction should automate issuance and rotation (Vault, cert-manager, or mesh CA). Keep CA keys offline or in HSM/KMS; distribute only ca.crt as the trust anchor.
Step 2 — Verify chain and identity before you wire apps
openssl verify -CAfile ca/ca.crt server/server.crt
openssl verify -CAfile ca/ca.crt client/client.crt
openssl x509 -in server/server.crt -noout -dates -ext subjectAltName
openssl x509 -in client/client.crt -noout -dates -ext subjectAltNameStep 3 — Terminate mTLS on NGINX (API or sidecar pattern)
server {
listen 8443 ssl;
server_name orders.internal;
ssl_certificate /etc/mtls/server.crt;
ssl_certificate_key /etc/mtls/server.key;
ssl_client_certificate /etc/mtls/ca.crt;
ssl_verify_client on;
ssl_verify_depth 2;
ssl_protocols TLSv1.2 TLSv1.3;
location / {
proxy_pass http://127.0.0.1:8080;
proxy_set_header X-Client-DN $ssl_client_s_dn;
proxy_set_header X-Client-Verify $ssl_client_verify;
}
}sudo nginx -t && sudo systemctl reload nginxStep 4 — Call with a real client cert (curl smoke test)
curl -v --http1.1 \
--cacert /etc/mtls/ca.crt \
--cert /etc/mtls/client.crt \
--key /etc/mtls/client.key \
https://orders.internal:8443/healthzExpect HTTP 200 and, on the server, SUCCESS for client verify. A failure without --cert/--key proves the server is rejecting anonymous TLS—exactly what you want for zero-trust east–west traffic.
Step 5 — Application-level clients (Node example)
// use tls options; never disable rejectUnauthorized in production
const fs = require('fs');
const https = require('https');
const agent = new https.Agent({
ca: fs.readFileSync('/etc/mtls/ca.crt'),
cert: fs.readFileSync('/etc/mtls/client.crt'),
key: fs.readFileSync('/etc/mtls/client.key'),
minVersion: 'TLSv1.2'
});
https.get('https://orders.internal:8443/healthz', { agent }, (res) => {
console.log('status', res.statusCode);
}).on('error', console.error);For multi-tenant SaaS and full-stack platforms that need identity-aware service meshes as they scale, Fixwebnode also delivers implementation work under Multi-Tenant SaaS Platform Development & Scalability | Melbourne and application builds via Custom Full-Stack Web Apps | Melbourne CBD Experts—still as a direct specialist engagement, not a marketplace bid.
Fix issue 1 — Incomplete CA chain
Symptom: unknown authority only on some services after you rotated the intermediate.
Step 1 — Dump what the peer actually presents
openssl s_client -connect orders.internal:8443 -showcerts -servername orders.internal </dev/null 2>/dev/null | openssl crl2pkcs7 -nocrl -certfile /dev/stdin 2>/dev/null | openssl pkcs7 -print_certs -nooutStep 2 — Confirm the local trust anchor matches the issuer
openssl x509 -in /etc/mtls/ca.crt -noout -subject -issuer -fingerprint -sha256
openssl x509 -in /etc/mtls/server.crt -noout -issuer -subjectStep 3 — Install the full chain the server should send (leaf + intermediate if used)
cat server/server.crt ca/intermediate.crt > /etc/mtls/server-fullchain.crt
# point ssl_certificate (or Kubernetes tls.crt) at server-fullchain.crt
sudo systemctl reload nginxStep 4 — Mount the same CA bundle on every client (ConfigMap/Secret or host path). Re-run the curl test with only --cacert pointing at that bundle.
When DIY is enough: single CA, few services, you control every trust store. Book Fixwebnode when rotation must be zero-downtime across many namespaces or when mesh and non-mesh trust domains collide.
Fix issue 2 — SAN / SPIFFE identity mismatch
Symptom: trust works in OpenSSL verify against the file, but live dial fails on name.
Step 1 — Compare dialled name to certificate SANs
getent hosts orders.internal
openssl x509 -in /etc/mtls/server.crt -noout -ext subjectAltNameStep 2 — Re-issue the leaf with every DNS name and URI clients use (cluster DNS, external name, SPIFFE ID). Avoid relying on CN alone; modern verifiers prefer SAN.
Step 3 — For Istio/Linkerd-style meshes, align workload identity
kubectl -n default get peerauthentication -o yaml
kubectl -n istio-system logs deploy/istiod | grep -i 'SPIFFE\|CSR\|certificate' | tail -n 50Set PeerAuthentication to STRICT only after every workload has a sidecar and correct ServiceAccount identity. A PERMISSIVE middle state is fine for migration; leaving it forever re-opens MITM on plain ports.
Step 4 — Pin verification in code to the expected identity (URI SAN or DNS), not to “any cert signed by CA,” if your threat model includes stolen leaves inside the trust domain.
Call a pro when custom resolvers, multi-cluster SPIFFE trust bundles, or tenant-specific hostnames make SAN matrices error-prone.
Fix issue 3 — Clock skew and short-lived certificates
Symptom: random handshake failures after cert-manager or Vault issues two-hour leaves.
Step 1 — Check time on both ends
timedatectl status
chronyc tracking 2>/dev/null || ntpq -p 2>/dev/null
date -u; ssh peer-host 'date -u'Step 2 — Inspect validity windows on the failing node
openssl x509 -in /etc/mtls/client.crt -noout -dates
openssl x509 -in /etc/mtls/server.crt -noout -datesStep 3 — Enforce NTP and restart the agent that mounts renewed secrets
sudo systemctl enable --now chronyd
# Kubernetes example: bounce pods so projected secrets remount
kubectl -n default rollout restart deploy/orders deploy/checkoutStep 4 — Add renewal headroom: renew at 60–70% of lifetime, and monitor NotAfter with a simple check in CI or a cron job that fails the build if production leaves expire inside seven days.
Book Fixwebnode if drift is fleet-wide, nodes disagree across regions, or automated renewals restart traffic at peak without dual-cert overlap.
Fix issue 4 — Client certificate not presented
Symptom: server TLS works from browsers or health checks that only do one-way TLS; internal auth still trusts source IP.
Step 1 — Force verify on the server (ssl_verify_client on or mesh STRICT). Reload and confirm anonymous curl fails.
curl -vk --cacert /etc/mtls/ca.crt https://orders.internal:8443/healthz
# expect handshake failure or 400/403 without client certStep 2 — Confirm the client process loads key + cert (file permissions 600, correct path inside the container).
ls -l /etc/mtls/
openssl rsa -in /etc/mtls/client.key -check -noout
openssl x509 -in /etc/mtls/client.crt -noout -purpose | headStep 3 — Map verified DN or SPIFFE ID to authorisation in the app (allowlist service identities), not only “TLS succeeded.”
Escalate when legacy libraries cannot supply client certs or when you must insert an mTLS gateway in front of apps you cannot recompile quickly.
Diagnostics checklist (remote-friendly)
- Capture a failing handshake:
openssl s_client -connect host:port -cert client.crt -key client.key -CAfile ca.crt -tlsextdebug - Application logs: NGINX
ssl_client_verify, Envoyresponse_flags, Javajavax.net.ssldebug only in a controlled window - Confirm no accidental plaintext listeners remain on the old port after cutover
- Rotate one service at a time; keep a rollback path that re-enables PERMISSIVE or dual stack briefly
sudo ss -lptn | grep -E ':80|:443|:8443'
sudo tail -n 100 /var/log/nginx/error.log
kubectl -n default get secret orders-mtls -o jsonpath='{.data.tls\.crt}' | base64 -d | openssl x509 -noout -dates -ext subjectAltNameWhen DIY is enough vs when to book Fixwebnode
DIY is enough when you own a small set of services, can restart them in a maintenance window, and the failures match the four patterns above. Stay hands-on through the OpenSSL verifies and one successful mutual curl before you declare victory.
Book Fixwebnode when you need a production cutover plan across many services, mixed mesh and bare-metal TLS, multi-tenant identity isolation, or a remote review of CA hierarchy and rotation without downtime. We work as a direct specialist provider for individuals, sole traders, and local operators across Australia—remote sessions focused on this man-in-the-middle network fix, not a board of competing freelancers.
Close the MITM gap—talk to Fixwebnode
Mutual TLS turns a breached node from a listening post into a dead end: peers refuse unauthenticated connections and payloads stay encrypted. If you want a second pair of specialist hands on certificates, mesh policy, or a staged STRICT rollout, start a conversation through https://fixwebnode.com.au/website-repair-australia. Bring your failing handshake logs and trust-store layout; we will help you finish an encrypted zero-trust pipeline your microservices can rely on.