Kubernetes Infrastructure: User Guide

The other pages in this section document what each screen shows; this one is about combining them to answer a specific question — why is this pod crashing, why won’t this pod schedule, why is this node unhealthy, why isn’t traffic reaching this service. Every scenario below uses only what’s already on the screen: a table column, a detail-pane tab, or the search bar/tags-and-labels panel.

General approach

A few habits apply to every scenario below:

  • Start from Status, not from a guess. The Status column on Pods and Nodes tables is the fastest way to find the object that’s actually broken, before you open anything. Use the search bar (pod_status=CrashLoopBackOff, for example) or the tags-and-labels panel to filter the table down to just the unhealthy objects.

  • Pivot through row links instead of re-filtering by hand. A Pod row links to its Cluster and Namespace; clicking either opens that object’s own detail pane scoped to the same time range, so you can move from "this pod" to "everything else in this namespace" in one click.

  • Read Events before Logs. The Events tab records what Kubernetes itself did or tried to do — scheduling decisions, restarts, failed probes, eviction warnings. It often names the actual cause (FailedScheduling, Unhealthy, BackOff) before you have to go looking for it in application output. Only then check Logs for what the application itself reported.

  • Cross-check YAML against Details. The Details tab is a curated summary; the YAML tab is the full manifest. When something’s misconfigured — a missing resource limit, a wrong image tag, an absent label a Service selector depends on — it’s visible in YAML even when it doesn’t have its own field in Details.

A pod is crashing or restarting repeatedly

  1. Open Pods and filter by status — search pod_status=CrashLoopBackOff, or check the Status column after grouping the table by it.

  2. Click the pod to open its detail pane. The Details tab’s Restarts count confirms how often this has happened, not just that it’s happening now.

  3. Check Events first — a BackOff or Unhealthy event usually names the immediate trigger (a failed liveness probe, an OOM kill, a failed image pull).

  4. Check Logs, scoped to the specific container if the pod runs more than one, for the application’s own error output at the moment of the crash — narrow the time range to just before the last restart.

  5. If the crash looks configuration-related rather than code-related (wrong environment variable, missing mount, probe pointed at the wrong port), check YAML for the container spec.

  6. Check Metrics for a memory usage line that climbs and cuts off right before each restart — that pattern points at an OOM kill even before Events confirms it.

A pod is stuck Pending

A Pending pod hasn’t failed — it hasn’t been scheduled yet, which is a different problem to diagnose.

  1. Open the pod’s detail pane and check Events first. FailedScheduling events explain why the scheduler is refusing to place it — insufficient CPU/memory on every node, an unsatisfied node selector or taint, or a volume that can’t be attached.

  2. If the event points at insufficient resources, open Nodes and compare CPU Allocable/Used and Memory Allocable/Used across the cluster’s nodes — a cluster that looks fine in aggregate can still have no single node with enough free capacity for this one pod’s request.

  3. If the event points at a volume, check the pod’s PersistentVolumeClaim — see A PersistentVolumeClaim won’t bind below.

  4. Check the pod’s YAML for its resource requests and any nodeSelector/affinity/tolerations blocks — a request that’s larger than any node’s capacity, or a selector that matches no node’s labels, produces exactly this symptom.

A node is NotReady or under pressure

  1. Open Nodes and filter or sort by Status to find nodes that aren’t Ready, and by CPU Used/Memory Used to find nodes running hot even while technically Ready.

  2. Open the node’s detail pane. Events records condition transitions — NodeNotReady, MemoryPressure, DiskPressure — with a timestamp for when the condition started.

  3. Check Metrics for the node’s own CPU, memory, network, and disk usage over time, to see whether the condition lines up with a genuine resource exhaustion or a transient spike.

  4. From here, pivot to the workloads affected: search Pods for kube_node=<node> to see every pod scheduled on the unhealthy node — pods on a NotReady node are the ones most likely to be evicted or to show connectivity problems next.

  5. If the node is schedulable but shouldn’t be receiving new pods while you investigate, its Details tab’s Schedulable field tells you whether it’s already cordoned.

A service isn’t reaching its pods

  1. Open the Service’s detail pane and check YAML for its selector — this is the set of labels it uses to find backend pods, and it’s the single most common cause of "the service exists but nothing’s behind it."

  2. Open the Metrics tab. Because a Service’s metrics are aggregated by the pods it fronts (see The resource detail pane), an empty or flat CPU/memory-by-pod chart is itself a signal that no pods currently match the selector.

  3. Open Pods, filter to the same namespace, and compare each candidate pod’s Labels (on its own Details tab) against the Service’s selector — a typo or a missing label here is the usual root cause.

  4. If the selector and labels match but traffic still isn’t arriving, check whether the request is coming through an Ingress or a Gateway/HTTPRoute pointed at this Service — open that object’s YAML and confirm its backend/rule references this Service’s name and port correctly.

A PersistentVolumeClaim won’t bind

  1. Open the PVC under Storage and check its Phase column/field — Pending means it hasn’t bound yet.

  2. Check Events on the PVC for provisioning failures (for example, a StorageClass with no available provisioner, or a capacity request larger than anything available).

  3. Check the PVC’s YAML for its requested storage class and access mode, and compare against the cluster’s StorageClasses table to confirm the class exists and its provisioner is what you expect.

  4. If the PVC references a specific PersistentVolume rather than relying on dynamic provisioning, open that PV and check its Phase and Reclaim Policy — a PV stuck in Released from a previous claim won’t bind to a new one without manual intervention.

Confirming (or ruling out) a bad deployment rollout

  1. Open the Deployment and compare Current, Desired, Up-to-date, and Available — a gap between Desired and Available right after a deploy means new pods aren’t becoming ready.

  2. Open the Deployment’s YAML to confirm which image/tag actually rolled out — useful when you need to confirm what shipped, not just that something shipped.

  3. Search Pods for the namespace and look for pods with high Restarts or a non-Running Status created around the same time as the rollout — the Age column makes the newest generation of pods easy to spot.

  4. Drill into one of the new, unhealthy pods and follow A pod is crashing or restarting repeatedly above.

Investigating a namespace after an alert

When an alert names a namespace rather than a specific object, start broad and narrow down:

  1. Open Namespaces, find the namespace, and check its CPU/Memory usage and Pods/Deployments/DaemonSets/StatefulSets/CronJobs counts for anything obviously out of the ordinary (a pod count much higher or lower than usual, for instance).

  2. Open the namespace’s detail pane and check Events for cluster-level problems affecting everything in it (quota exceeded, admission webhook failures) before assuming the issue is one specific pod.

  3. Search Pods for kube_namespace=<namespace> and sort by Restarts or group by Status to find the specific unhealthy objects, then follow the relevant scenario above for each one.