Logs: User Guide
The other pages in this section document what each screen does; this one is about combining them to actually debug something — narrowing millions of lines down to the handful that explain a problem, and making that investigation reusable instead of starting over next time.
General approach
-
Filter before you search text. A field filter (
level=error,kube_namespace=payments) is cheap because the Logs Store already has labels and facets indexed; a term or regex search has to touch the message text itself — see How a search reaches the data. Narrow with fields first, then add text search only for what isn’t already a field. -
Reach for Fingerprints before you reach for "scroll and read." When a source is noisy, clustering by message template (Log Fingerprints) tells you what kinds of lines it’s producing in one screen, instead of you reading hundreds of near-duplicates to notice the same handful repeating.
-
Pivot on shared labels, not on log-line intuition. A log’s Kubernetes and Cloud fields are the same fields shown on the pod, node, or service that produced it elsewhere in Kloudfuse — see Facets, labels, and tags — and why the distinction matters for search. Use them to jump from "this log line" to "everything else from this pod" without retyping a filter by hand.
Investigating an error spike
-
Start broad:
level=error, scoped to the time window the spike happened in, in the Timeseries view groupedby source— this immediately shows which source(s) are producing the spike rather than just that error volume went up somewhere. -
Switch to Fingerprints with the same filter. Rather than reading individual error lines, this clusters them by message template — a spike that’s actually five occurrences of one new error looks very different from five hundred occurrences of one existing, already-understood error, and Fingerprints tells them apart immediately.
-
Once you’ve identified the fingerprint(s) driving the spike, switch back to Logs and search that fingerprint’s distinguishing text to see full example lines, including the facets extracted from them (timestamps, IDs, durations — whatever the message contains).
-
Check whether the spike correlates with a deploy or a config change: group the Timeseries view by a version or revision label if one is present on your logs, or cross-reference the timing against a recent deployment rollout in Kubernetes Infrastructure: Pods if the source runs on Kubernetes.
Tracing what happened to one request
-
If you have any distinguishing value for the request — a request ID, a user ID, an IP, a trace ID — search for it as a field filter if it’s already extracted as a facet (
@request_id==…), or as a substring search ("req-abc123") if it’s only ever appeared in message text. -
Open a matching log’s detail pane and use Show in context to see the surrounding lines from the same source — this is often enough to see what happened immediately before and after without constructing a wider time-range search.
-
If the request touched multiple services, repeat the same ID search with
kube_namespaceorsourceremoved from your filter, so you’re searching across every service rather than just the one you started in. -
If the log lines don’t fully explain the failure and the service is instrumented for tracing, pivot to the equivalent trace in APM Traces using the same request or trace identifier — the log gives you the text; the trace gives you the timing and the call graph.
Understanding what a noisy or unfamiliar source actually logs
-
Filter to the source alone, with no other conditions, over a representative time window (an hour is usually enough).
-
Switch to Fingerprints and group by nothing else — this lists every distinct message template the source produces, ranked by occurrence count. This is the fastest way to learn "what does this thing normally say" before you’ve read a single full line.
-
Sort by count ascending instead of descending to find the rare fingerprints — these are disproportionately likely to be the interesting ones (a one-off warning, a rare error path) compared to the high-volume routine ones.
-
If a fingerprint’s facets aren’t precise enough to be useful (everything lumped into one generic placeholder), that’s a sign to define a custom grammar pattern for this source rather than relying on automatic extraction.
Cleaning up a noisy or expensive label
-
Open Logs Cardinality and sort by Value count to find labels with unexpectedly many distinct values.
-
For each one, ask whether the values are actually useful to filter or group by — a label whose values are effectively unique per log line (a generated ID, a timestamp fragment) is expensive to index and rarely worth keeping as a label rather than leaving it in the message or as a facet.
-
Drop or rename the label at ingestion with a Relabel rule, rather than trying to filter it out after the fact in every query.
Turning a one-off investigation into something reusable
Once a search proves useful, don’t rebuild it next time:
-
Reuse it yourself later — save it as a saved query.
-
Get notified automatically going forward — turn the same query into a log alert if it should page someone the moment it’s true again, or a Scheduled Search if you want a recurring digest instead.
-
Make it fast at any time range — if it’s an aggregation you’ll run often over long windows (volume by source, error rate by service), define it as a Scheduled View instead of re-aggregating raw logs every time.