Log Parsing Configuration
Kloudfuse extracts structure from every log line automatically — a fingerprint, a set of facets, and a timestamp — using the parsing pipeline described in How a log line reaches the Logs Store. This page is the configuration reference for that pipeline: how to shape what it extracts, when the automatic heuristics aren’t precise enough. All of it is configured under kf_parsing_config in the logs-parser Helm values.
Remap (deprecated)
This YAML-configured remap stage is being phased out in favor of Logs Remap — a newer engine with the same job (normalizing each agent’s raw payload into message/timestamp/source/labels) configured from an editable UI page instead of Helm values, with a live preview against real traffic. See Logs Remap and Remap Architecture. New deployments should configure normalization there; this stage remains only for deployments still running with global.logs.remap.enabled: false — see Enabling the remap engine for what that setting controls and how to tell which path your deployment is on.
|
Remap is the pipeline’s first stage — it maps fields from the incoming payload (JSON, msgpack, or proto, depending on your agent) onto Kloudfuse’s internal fields: message, timestamp, source, and initial labels/facets.
logs-parser:
kf_parsing_config:
config: |-
- remap:
args:
kf_source:
- "$.logSource"
kf_msg:
- "$.logMessage"
conditions:
- matcher: "__kf_agent"
value: "fluent-bit"
op: "=="
__kf_agent is a reserved matcher for the sending agent, and supports datadog, fluent-bit, fluentd, kinesis, gcp, otlp, and filebeat. All Remap fields must be specified in JSONPath notation.
Relabel
Relabel operates on the labels Remap (or Logs Remap) produced — add, drop, rewrite, or derive a label. It’s loosely inspired by Prometheus relabeling (regex against one or more source labels, an action, a target), but the action set below is this pipeline’s own, not Prometheus’s.
This relabel function is specific to the logs-parser pipeline and only affects logs. Metrics, Events, and Traces have a separate, ingester-level relabeling mechanism that genuinely does follow the Prometheus relabel_config format — see Relabel Rules.
|
Actions
action selects what the function does with sourceLabels (a comma-separated list — facets prefixed @, labels prefixed # or bare):
- replace
-
Regex-matches
sourceLabels(joined byseparator) againstregex, and if it matches, writesreplacementtotargetLabel.targetLabelprefixed@writes a facet instead of a label. - drop
-
Discards the log line entirely if
sourceLabelsmatchesregex. - keep
-
Discards the log line entirely if
sourceLabelsdoes not matchregex. - lowercase
-
Writes the lower-cased value of
sourceLabelstotargetLabel. - uppercase
-
Writes the upper-cased value of
sourceLabelstotargetLabel. - label_map
-
Renames every label whose name matches
regex, replacing the matched part withreplacement. Unlike the other actions, this operates on label names, not values, and ignoressourceLabels. - label_keep
-
Removes every label whose name does not match
regex. IgnoressourceLabels. - label_drop
-
Removes every label whose name matches
regex. IgnoressourceLabels. - facet_to_label_map
-
Copies the value of a facet in
sourceLabelstotargetLabelas a label — withregex/replacement, transforms the value on the way; with noregex, copies it as-is. This is the action to use for Transform's "promote a facet to a label" use case.
- relabel:
args:
- action: "replace"
- regex: ".*"
- replacement: "production"
- targetLabel: "env"
- relabel:
args:
- action: "drop"
- sourceLabels: "@path"
- regex: "/healthz"
Functions and conditions
Every pipeline function — relabel, transform, parser (dissect/grok), and the others listed under Other pipeline functions — is a list item under config, with its arguments under args and, optionally, a list of conditions that determine whether it runs for a given log line. remap is the one exception: it always runs first, ahead of everything else in the list. Every other function runs in the order you list it, which is what actually determines whether, say, a relabel rule sees a facet a later parser step hasn’t extracted yet. A function with no conditions always runs. Each condition is a matcher / value / op triplet:
- matcher
-
A label (
#name), a facet (@name), the message field (%kf_msg— the only supported field), a JSONPath from the incoming payload (remap only), or the reserved__kf_agent(remap only). - value
-
A literal string, or a regex depending on
op. - op
-
One of
==,!=,=~(regex match),!~(regex not match),contains,startsWith,endsWith, orin(value is one of a comma-separated list) — each also has a negated form,not contains,not startsWith,not endsWith, andnot in.opis case-insensitive.
logs-parser:
kf_parsing_config:
configPath: "/conf"
config: |-
- <FUNC_NAME>:
args:
- <FUNC_ARGS>
conditions:
- matcher: <LABEL|FACET|FIELD|JSON_PATH|__kf_agent>
value: "<EXPECTED_VALUE>"
op: "<OP_VALUE>"
Grammar: dissect and grok patterns
Kloudfuse auto-detects facets and a timestamp from the message using a heuristic, which won’t always be precise. Define a custom pattern to extract fields exactly instead.
Dissect patterns
A dissect pattern is a text tokenizer: sections like %{fieldname} separated by literal delimiter text. Prefixes on the field name change its behavior — %{?a} matches without capturing, %{+b} appends to a previous capture, %{&c} references an earlier capture. Use bracket notation (%{[field][subfield]}) to generate a nested field name rather than one containing a literal dot. Test patterns with a dissect debugger before deploying them.
- parser:
dissect:
args:
- tokenizer: '%{timestamp} %{level} [LLRealtimeSegmentDataManager_%{segment_name}]'
conditions:
- matcher: "%kf_msg"
value: "LLRealtimeSegmentDataManager_"
op: "contains"
Grok patterns
A grok pattern is a set of named regexes — see the Grok filter plugin documentation for the pattern syntax. Test with a grok debugger before deploying.
Apply one or more patterns to a field with args: patterns (a list — combined together) and optionally field (defaults to the message):
- parser:
grok:
args:
- patterns:
- '%{TIMESTAMP_ISO8601:ts} %{LOGLEVEL:level} %{GREEDYDATA:msg}'
conditions:
- matcher: "#source"
value: "nginx"
op: "=="
Reusable pattern definitions go under parser_patterns, and can be referenced by name elsewhere in the config:
parser_patterns:
- dissect:
timestamp_pat: "<REGEX>"
- grok:
NGINX_HOST: (?:%{IP:destination_ip}|%{NGINX_NOTSEPARATOR:destination_domain})(:%{NUMBER:destination_port})?
Worked example: deriving a facet with a tokenizer
Given this log line:
10.12.0.35 - - [26/May/2021:18:59:10 +0000] "GET /unavailable HTTP/1.1" 503 21 "-" "hey/0.0.1"
this dissect tokenizer:
'%{sourceIp} - - [%{timestamp}] "%{requestMethod} %{uri} %{_}" %{responseCode} %{contentLength}'
produces the facets sourceIp: 10.12.0.35, requestMethod: GET, responseCode: 503, and contentLength: 21. Scope a tokenizer to one source with a conditions matcher on #source, so it only runs against the logs it’s meant for:
logs-parser:
kf_parsing_config:
config: |-
- parser:
dissect:
args:
- tokenizer: '%{sourceIp} - - [%{timestamp}] "%{requestMethod} %{uri} %{_}" %{responseCode} %{contentLength}'
conditions:
- matcher: "#source"
value: "nginx"
op: "=="
Add another - parser: dissect: … item, with its own conditions, for a second source — each function in the config list is independently scoped by its own conditions block; nothing groups them by source implicitly.
JSON logs
Kloudfuse detects JSON-formatted messages automatically and extracts every field as a facet, flattening nested objects with _ — {"location": {"city": "SF"}} becomes the facet location_city. An array field like "aliases": ["johndoe", "johnny"] is extracted as a single string facet.
A single log line can produce at most 50 facets by default; extraction fails silently beyond that limit. Raise it with skipAutoFacet:
logs-parser:
kf_parsing_config:
config: |-
...
- skipAutoFacet:
args:
maxFacetsCount: 100
Once ingested, query the extracted facets from the Logs Explorer the same way as any other facet — see Logs Search Syntax.
Transform and fingerprinting
transform runs the exact same engine and action set as Relabel above — it’s the same function, just conventionally placed later in your config list, after the functions that extract the facets it promotes. The most common use is facet_to_label_map, promoting a facet extracted in Grammar (above) into a label — for example promoting a facet eventSource to a label source:
- transform:
args:
- action: "facet_to_label_map"
- sourceLabels: "@eventSource"
- targetLabel: "source"
conditions:
- matcher: "#source"
op: "=="
value: "awsLogSource"
Any other action from Relabel’s action table — drop, keep, label_map, and so on — works identically here; the name transform versus relabel reflects where you’ve placed it in the list, not a difference in capability.
Somewhere in the pipeline, the internal Kf-Parse stage also generates the log’s fingerprint — splitting the message into its static template and dynamic values, so that ts=2023-02-01T22:51:33Z caller=logging.go:29 method=Authorise result=false took=9.775µs becomes the fingerprint ts=<v_0> caller=<v_1> method=<v_2> result=<v_3> took=<v_4> plus the facets ts, caller, method, result, and took. Storing the template once and the variable values separately is what keeps Kloudfuse’s highly repetitive log traffic both fully indexed and storage-efficient — see Log Fingerprints for where this shows up in the Explorer. Kf-Parse itself isn’t configurable.
See also Log Timestamp Handling for how Kloudfuse selects and validates a log’s timestamp specifically.
Other pipeline functions
Beyond remap, relabel/transform, and dissect/grok, the pipeline supports these additional functions — each a list item under config, exactly like the examples above:
| Function | What it does |
|---|---|
|
Adds a facet with a fixed value. Args: |
|
Removes one or more named facets. Args: |
|
Renames a facet. Args: |
|
Moves a label’s value into a facet of the same or a different name. Args: |
|
Forces the log’s level to a fixed value, overriding whatever was detected. Args: |
|
Parses a facet as the log’s timestamp using one or more explicit formats, instead of relying on automatic timestamp detection. Args: |
|
Tunes how far into the message the automatic timestamp-detection heuristic looks. Args: |
|
Restricts extraction to only the named facets, discarding everything else auto-detection would otherwise extract. Args: |
|
Parses the message as logfmt (space-separated |
|
Unconditionally drops the log line — pair it with |
|
Controls this log’s archival behavior — which archive it’s written to, and whether it’s indexed for live search as well. See Archiving Logs and Hydration. |
A further set of low-level tokenization primitives — resetPos, truncate, tokenize, rebuild, tokenizedKvExtract, tokenizedTsExtract, dissectKVOptimize, and tokenizedNewr — exist for advanced custom parsing scenarios beyond what dissect and grok cover directly. They’re implementation details rather than something most configurations need; contact your Kloudfuse representative if dissect or grok don’t cover a format you’re trying to parse.