This is unreleased documentation for Network Enforcer 0.3-dev.

Troubleshooting

This document helps diagnose common Kubewarden Network Enforcer failures. Failure modes are provider-specific: whether Goldmane is up, whether Hubble Relay is reachable, or whether a namespace is ambient-enrolled, cannot be inferred from the CRD alone.

Examples below assume the quickstart release name network-enforcer in namespace network-enforcer. Adjust names if you installed differently.

Enable verbose logs

To enable more verbose logs via Helm:

helm upgrade network-enforcer <chart> \
  --namespace network-enforcer \
  --set controller.logLevel=debug \
  --reuse-values

controller.logLevel accepts debug, info, warn, or error. The chart passes the value as --log-level on the controller.

Then follow the controller logs:

kubectl logs -n network-enforcer \
  -l app.kubernetes.io/component=controller -c manager -f

No proposals are being created

Common checks

  • Controller pod is Running.

  • controller.provider.name matches the data plane in the cluster (istio, calico, or cilium).

  • Workloads are a supported owner kind: Deployment, StatefulSet, or DaemonSet.

  • Traffic happened after the controller started. Learning does not reconstruct historical flows; restart workloads or generate traffic while Network Enforcer is running.

Istio

  1. Namespace labeled istio.io/dataplane-mode=ambient before the pods were created. istio-cni decides redirection at pod creation time; labeling later does not enroll existing pods.

  2. ztunnel and istiod flags required for usable access logs:

    • ztunnel: AUTHZ_POLICY_INFO_LOGGING=true and logAsJson=true

    • istiod: AMBIENT_ENABLE_DRY_RUN_AUTHORIZATION_POLICY=true (also required for monitor-mode dry-run evaluation)

  3. DaemonSet <fullname>-istio-fluent-bit is running. It tails /var/log/containers/ztunnel.log and ships OTLP/gRPC to the controller Service <fullname>-istio-otlp on port 4317.

  4. Confirm fluent-bit pods and the OTLP Service:

    kubectl get ds,pods -n network-enforcer -l app.kubernetes.io/component=istio-fluent-bit
    kubectl get svc -n network-enforcer network-enforcer-istio-otlp
UDP never goes through ztunnel, so it never produces proposals. Istio learning creates ingress proposals on the destination only.

Do not confuse the Istio ingestion port (<fullname>-istio-otlp:4317) with the violation-export collector (<fullname>-otel-collector:4317). They are separate OTLP hops.

Calico

  1. Goldmane is enabled in the Tigera Installation (goldmane.enabled=true).

  2. Secret goldmane-key-pair and ConfigMap goldmane-ca-bundle exist in calico-system.

  3. Default endpoint goldmane.calico-system.svc:7443 is reachable from the controller (override with controller.provider.calico.endpoint).

    kubectl get secret -n calico-system goldmane-key-pair
    kubectl get configmap -n calico-system goldmane-ca-bundle
    kubectl get deploy -n calico-system -l k8s-app=goldmane

Cilium

  1. Hubble and Hubble Relay are enabled (hubble.enabled=true, hubble.relay.enabled=true).

  2. Secret hubble-relay-client-certs exists in kube-system.

  3. Default endpoint hubble-relay.kube-system.svc:443 is reachable (override with controller.provider.cilium.endpoint).

    kubectl get secret -n kube-system hubble-relay-client-certs
    kubectl get svc -n kube-system hubble-relay
    kubectl get pods -n kube-system -l k8s-app=hubble-relay

Provider TLS handshake

Read the active mode from the controller arguments:

kubectl get deploy -n network-enforcer \
  -l app.kubernetes.io/component=controller \
  -o jsonpath='{range .items[0].spec.template.spec.containers[0].args}{@}{"\n"}{end}' \
  | grep provider-

--provider-tls-mode is issuer, existingSecret, or insecure. The same output shows --provider-endpoint and the certificate flags. issuer and a same-namespace Secret mount volume provider-tls at /etc/provider/certs. When existingSecret.namespace is set, the controller reads the Secret through the API.

Raise controller.logLevel to debug before you follow a handshake. The commands are in Enable verbose logs.

Istio

The controller is the TLS server. fluent-bit is the client. The default mode is issuer.

  1. If the controller pod or a fluent-bit pod stays Pending, describe the pod. A missing cert-manager-csi-driver fails the provider-tls volume.

  2. If the controller log says failed to load TLS credentials for OTLP logs server, /etc/provider/certs is missing tls.crt, tls.key, or ca.crt.

  3. If the controller log says OTLP Istio scraper running in insecure mode, TLS is off. fluent-bit uses the same mode. A mismatch fails the client handshake.

  4. If the controller is running and proposals stay empty, read the fluent-bit logs:

    kubectl logs -n network-enforcer \
      -l app.kubernetes.io/component=istio-fluent-bit --tail=100

    fluent-bit sets Tls.verify_hostname On unless the mode is insecure. A certificate name that does not match Service <fullname>-istio-otlp shows up in those logs.

Calico

Goldmane requires mutual TLS. The chart rejects mode insecure.

  1. Make sure that Secret goldmane-key-pair and ConfigMap goldmane-ca-bundle exist in calico-system. The Secret must contain tls.crt and tls.key. The ConfigMap must contain key tigera-ca-bundle.crt.

  2. Make sure that Role <fullname>-provider-tls in calico-system allows get on those objects. The log failed to get Secret means the Secret is missing, or the Role does not allow get. The log failed to get ConfigMap means the same for the ConfigMap.

  3. The log secret …​ is missing key means tls.crt or tls.key is absent. The log configmap …​ is missing key means the CA key is absent.

  4. A successful dial logs Using TLS credentials for Goldmane connection and a serverName.

  5. If a later log names a certificate or a handshake, the material loaded and Goldmane rejected it. An empty controller.provider.calico.tls.serverName uses the endpoint host goldmane.calico-system.svc. Set serverName when the Goldmane certificate uses a different name.

If existingSecret.namespace is empty, the Secret is mounted from the release namespace. --provider-tls-cert-dir=/etc/provider/certs is present and --provider-tls-cert-secret is absent. The mounted Secret must contain tls.crt, tls.key, and ca.crt.

Cilium

The default mode is existingSecret. The default endpoint port is 443. The default server name is ui.hubble-relay.cilium.io.

  1. Make sure that Secret hubble-relay-client-certs in kube-system contains ca.crt, tls.crt, and tls.key.

  2. The log failed to obtain TLS material for Hubble Relay means the Secret read failed. The log failed to load TLS credentials for Hubble Relay means the PEM data is not usable.

  3. A successful dial logs Using TLS credentials for Hubble Relay connection and a serverName.

  4. If that line is followed by a handshake error, the server name does not match the Relay certificate. Relay certificates use a name under hubble-relay.cilium.io.

  5. The log Connecting to Hubble Relay without TLS means the mode is insecure. Relay must listen without TLS. Clear controller.provider.cilium.tls.serverName. The chart automatically overwrites only the default endpoint hubble-relay.kube-system.svc:443 to port 80. If --provider-endpoint still uses port 443, set the endpoint to port 80.

Flow dumper

The flow dumper is an optional debug tool, disabled by default.

Features

Dump ingested flows

Each scraper (Istio OTLP logs, Goldmane, Hubble) writes marshaled records into an in-memory ring buffer. GET /flow drains that buffer and returns newline-delimited JSON (application/jsonl).

  • Empty body: nothing has been ingested since the last drain.

  • JSONL lines: ingestion is alive; inspect the payload for identity, ports, and direction.

GET /flow is destructive: each request drains the buffer (newest records first). If nobody is reading, the buffer silently drops when full.

Enable

helm upgrade network-enforcer <chart> \
  --namespace network-enforcer \
  --set controller.flowDumper.enabled=true \
  --reuse-values

Helm values:

Value Default Role

controller.flowDumper.enabled

false

Turns the HTTP dumper on and creates the Service.

controller.flowDumper.port

9080

Container and Service port.

controller.flowDumper.bufferSize

10000

In-memory ring buffer capacity.

The chart creates Service <fullname>-flow-dumper (network-enforcer-flow-dumper with the quickstart release name), type NodePort (node port 30080), selecting the controller pods.

Enabling the dumper restarts the controller (new args, container port, and Service).

Worked example

kubectl -n network-enforcer port-forward \
  svc/network-enforcer-flow-dumper 9080:9080

Generate traffic against the workloads you expect to learn, then:

curl -sS http://127.0.0.1:9080/flow

You should see one JSON object per line. A second immediate curl often returns empty until more flows arrive.

Violation pipeline debugging

Violations land on WorkloadNetworkPolicy status and, when telemetry is enabled, are also exported as OpenTelemetry logs.

Tail the OpenTelemetry collector

With telemetry.collectorStrategy=default, the release deploys <fullname>-otel-collector. Its logs pipeline uses the debug exporter (default verbosity normal).

For richer dumps while troubleshooting, set the exporter to detailed by editing the collector ConfigMap, then restart the collector pod:

kubectl -n network-enforcer edit configmap network-enforcer-otel-collector
# under exporters.debug, set: verbosity: detailed

kubectl -n network-enforcer rollout restart deploy/network-enforcer-otel-collector

Then follow the collector logs:

kubectl logs -n network-enforcer \
  -l app.kubernetes.io/component=otel-collector -f

Look for policy_violation_observed and policy_violation_acknowledged.

With telemetry.collectorStrategy=none or external, this in-cluster collector is not deployed. Status still records violations when export is off.

Scrape Prometheus metrics on :9090

The collector exposes Prometheus metrics on port 9090. Port-forward the collector Service and scrape:

kubectl -n network-enforcer port-forward \
  svc/network-enforcer-otel-collector 9090:9090

curl -sS http://127.0.0.1:9090/metrics | grep network_enforcer_policy_denies

The count connector feeds network_enforcer_policy_denies. Empty metrics are expected when the default collector is not installed.

The controller’s own metrics bind is HTTPS :8443 with auth; that is a different endpoint.

wnpStatusUpdateInterval latency

controller.wnpStatusUpdateInterval (default 30s) is how often the controller drains buffered violation observations into WorkloadNetworkPolicy status. Status is not patched per flow.

If a violation you just caused is missing from status.violations, wait at least one interval (and prefer watching with kubectl get wnp -w or a second kubectl get after ~30s) before reporting it as lost. The quickstart lowers this to 3s only to make demos responsive; production rarely needs to change the default.

Understand status fields

WorkloadNetworkPolicy.status

Field Meaning

status.observedGeneration

Last generation the status sync observed.

status.violationCount

Count of violation records for this policy, including trimmed or cleared entries. May be temporarily outdated until the next status sync.

status.activeViolationCount

Number of currently active (non-cleared, non-acknowledged) records.

status.violations

Most recent active records (cap 100).

status.acknowledgedViolations

Most recent acknowledged records (cap 100).

Acknowledge an expected violation by annotating the policy with networkenforcer.kubewarden.io/acknowledge-<id>, where <id> is the record id.

kubectl get wnp <name> -n <namespace> -o yaml
kubectl describe wnp <name> -n <namespace>

WorkloadNetworkPolicyProposal.status.conditions

The CRD defines status.conditions, but the controller does not write them today. An empty or absent conditions list is normal.

Promotion is label-driven, not condition-driven:

  • Set networkenforcer.kubewarden.io/promote=monitor or protect on the proposal.

  • The proposal reconciler creates a same-named WorkloadNetworkPolicy, sets networkenforcer.kubewarden.io/promoted-from: <proposal-name>, and deletes the proposal.

kubectl get wnpp,wnp -n <namespace>
kubectl get wnpp <name> -n <namespace> -o yaml

Enforcement mismatch

Inspect the native object the reconciler creates for the WorkloadNetworkPolicy (an Istio AuthorizationPolicy, or a Kubernetes NetworkPolicy on Calico and Cilium), then check other policies in the same namespace that may also allow or deny the traffic.

Traffic still flows

  1. Check the policy’s spec.mode, monitor mode never blocks traffic.

    • Istio: the controller still creates an AuthorizationPolicy annotated istio.io/dry-run=true.

    • Calico / Cilium: the controller will not create NetworkPolicy in monitor mode. If you switch back from protect, it will delete the NetworkPolicy.

  2. In protect mode, check that the native object exists and is owned by the WorkloadNetworkPolicy:

    # Istio
    kubectl get authorizationpolicy <wnp-name> -n <namespace> -o yaml
    
    # Calico / Cilium
    kubectl get networkpolicy <wnp-name> -n <namespace> -o yaml
  3. If a same-named NetworkPolicy or AuthorizationPolicy already existed and is not controlled by the WNP, the reconciler refuses to adopt it and logs refusing to manage existing … not controlled by a WorkloadNetworkPolicy. Rename or delete the conflicting object manually.

  4. Confirm the selector on the generated policy matches the workload pods.

Everything is blocked

  1. Pre-existing Kubernetes NetworkPolicy objects in the namespace AND with Network Enforcer’s policy. A broad default-deny plus a narrow learned allow list will drop unexpected peers.

  2. Extra Istio DENY (or restrictive ALLOW) AuthorizationPolicy objects.

  3. The learned allow list is too narrow because learning only saw traffic after install. Generate the missing flows in learn/monitor, or widen the policy.

  4. On Istio, a namespace that is not ambient-enrolled will not apply ztunnel authorization the way the quickstart expects.

CEL rejection messages

API server validation on PolicyBackendSpec emits these messages:

backend must match exactly one populated backend spec
kubernetes.podSelector cannot be empty: it must define at least one between matchLabel or matchExpression
istio.selector cannot be empty: it must define at least one matchLabel or matchExpression
backend is immutable

The chart also installs a ValidatingAdmissionPolicy that rejects an invalid promote label:

networkenforcer.kubewarden.io/promote must be "monitor" or "protect"