Log Parsing Configuration

Kloudfuse extracts structure from every log line automatically — a fingerprint, a set of facets, and a timestamp — using the parsing pipeline described in How a log line reaches the Logs Store. This page is the configuration reference for that pipeline: how to shape what it extracts, when the automatic heuristics aren’t precise enough. All of it is configured under kf_parsing_config in the logs-parser Helm values.

Remap (deprecated)

This YAML-configured remap stage is being phased out in favor of Logs Remap — a newer engine with the same job (normalizing each agent’s raw payload into message/timestamp/source/labels) configured from an editable UI page instead of Helm values, with a live preview against real traffic. See Logs Remap and Remap Architecture. New deployments should configure normalization there; this stage remains only for deployments still running with global.logs.remap.enabled: false — see Enabling the remap engine for what that setting controls and how to tell which path your deployment is on.

Remap is the pipeline’s first stage — it maps fields from the incoming payload (JSON, msgpack, or proto, depending on your agent) onto Kloudfuse’s internal fields: message, timestamp, source, and initial labels/facets.

logs-parser:
  kf_parsing_config:
    config: |-
      - remap:
          args:
            kf_source:
              - "$.logSource"
            kf_msg:
              - "$.logMessage"
          conditions:
            - matcher: "__kf_agent"
              value: "fluent-bit"
              op: "=="
yaml

__kf_agent is a reserved matcher for the sending agent, and supports datadog, fluent-bit, fluentd, kinesis, gcp, otlp, and filebeat. All Remap fields must be specified in JSONPath notation.

Relabel

Relabel operates on the labels Remap (or Logs Remap) produced — add, drop, rewrite, or derive a label. It’s loosely inspired by Prometheus relabeling (regex against one or more source labels, an action, a target), but the action set below is this pipeline’s own, not Prometheus’s.

This relabel function is specific to the logs-parser pipeline and only affects logs. Metrics, Events, and Traces have a separate, ingester-level relabeling mechanism that genuinely does follow the Prometheus relabel_config format — see Relabel Rules.

Actions

action selects what the function does with sourceLabels (a comma-separated list — facets prefixed @, labels prefixed # or bare):

replace

Regex-matches sourceLabels (joined by separator) against regex, and if it matches, writes replacement to targetLabel. targetLabel prefixed @ writes a facet instead of a label.

drop

Discards the log line entirely if sourceLabels matches regex.

keep

Discards the log line entirely if sourceLabels does not match regex.

lowercase

Writes the lower-cased value of sourceLabels to targetLabel.

uppercase

Writes the upper-cased value of sourceLabels to targetLabel.

label_map

Renames every label whose name matches regex, replacing the matched part with replacement. Unlike the other actions, this operates on label names, not values, and ignores sourceLabels.

label_keep

Removes every label whose name does not match regex. Ignores sourceLabels.

label_drop

Removes every label whose name matches regex. Ignores sourceLabels.

facet_to_label_map

Copies the value of a facet in sourceLabels to targetLabel as a label — with regex/replacement, transforms the value on the way; with no regex, copies it as-is. This is the action to use for Transform's "promote a facet to a label" use case.

Add a label env with value "production"
- relabel:
    args:
      - action: "replace"
      - regex: ".*"
      - replacement: "production"
      - targetLabel: "env"
yaml
Drop log lines from a noisy health-check endpoint
- relabel:
    args:
      - action: "drop"
      - sourceLabels: "@path"
      - regex: "/healthz"
yaml

Functions and conditions

Every pipeline function — relabel, transform, parser (dissect/grok), and the others listed under Other pipeline functions — is a list item under config, with its arguments under args and, optionally, a list of conditions that determine whether it runs for a given log line. remap is the one exception: it always runs first, ahead of everything else in the list. Every other function runs in the order you list it, which is what actually determines whether, say, a relabel rule sees a facet a later parser step hasn’t extracted yet. A function with no conditions always runs. Each condition is a matcher / value / op triplet:

matcher

A label (#name), a facet (@name), the message field (%kf_msg — the only supported field), a JSONPath from the incoming payload (remap only), or the reserved __kf_agent (remap only).

value

A literal string, or a regex depending on op.

op

One of ==, !=, =~ (regex match), !~ (regex not match), contains, startsWith, endsWith, or in (value is one of a comma-separated list) — each also has a negated form, not contains, not startsWith, not endsWith, and not in. op is case-insensitive.

logs-parser:
  kf_parsing_config:
    configPath: "/conf"
    config: |-
      - <FUNC_NAME>:
        args:
          - <FUNC_ARGS>
        conditions:
          - matcher: <LABEL|FACET|FIELD|JSON_PATH|__kf_agent>
            value: "<EXPECTED_VALUE>"
            op: "<OP_VALUE>"
yaml

Grammar: dissect and grok patterns

Kloudfuse auto-detects facets and a timestamp from the message using a heuristic, which won’t always be precise. Define a custom pattern to extract fields exactly instead.

Dissect patterns

A dissect pattern is a text tokenizer: sections like %{fieldname} separated by literal delimiter text. Prefixes on the field name change its behavior — %{?a} matches without capturing, %{+b} appends to a previous capture, %{&c} references an earlier capture. Use bracket notation (%{[field][subfield]}) to generate a nested field name rather than one containing a literal dot. Test patterns with a dissect debugger before deploying them.

- parser:
    dissect:
      args:
        - tokenizer: '%{timestamp} %{level} [LLRealtimeSegmentDataManager_%{segment_name}]'
      conditions:
        - matcher: "%kf_msg"
          value: "LLRealtimeSegmentDataManager_"
          op: "contains"
yaml

Grok patterns

A grok pattern is a set of named regexes — see the Grok filter plugin documentation for the pattern syntax. Test with a grok debugger before deploying.

Apply one or more patterns to a field with args: patterns (a list — combined together) and optionally field (defaults to the message):

- parser:
    grok:
      args:
        - patterns:
            - '%{TIMESTAMP_ISO8601:ts} %{LOGLEVEL:level} %{GREEDYDATA:msg}'
      conditions:
        - matcher: "#source"
          value: "nginx"
          op: "=="
yaml

Reusable pattern definitions go under parser_patterns, and can be referenced by name elsewhere in the config:

parser_patterns:
  - dissect:
      timestamp_pat: "<REGEX>"
  - grok:
      NGINX_HOST: (?:%{IP:destination_ip}|%{NGINX_NOTSEPARATOR:destination_domain})(:%{NUMBER:destination_port})?
yaml

Worked example: deriving a facet with a tokenizer

Given this log line:

10.12.0.35 - - [26/May/2021:18:59:10 +0000] "GET /unavailable HTTP/1.1" 503 21 "-" "hey/0.0.1"

this dissect tokenizer:

'%{sourceIp} - - [%{timestamp}] "%{requestMethod} %{uri} %{_}" %{responseCode} %{contentLength}'

produces the facets sourceIp: 10.12.0.35, requestMethod: GET, responseCode: 503, and contentLength: 21. Scope a tokenizer to one source with a conditions matcher on #source, so it only runs against the logs it’s meant for:

logs-parser:
  kf_parsing_config:
    config: |-
      - parser:
          dissect:
            args:
              - tokenizer: '%{sourceIp} - - [%{timestamp}] "%{requestMethod} %{uri} %{_}" %{responseCode} %{contentLength}'
            conditions:
              - matcher: "#source"
                value: "nginx"
                op: "=="
yaml

Add another - parser: dissect: …​ item, with its own conditions, for a second source — each function in the config list is independently scoped by its own conditions block; nothing groups them by source implicitly.

JSON logs

Kloudfuse detects JSON-formatted messages automatically and extracts every field as a facet, flattening nested objects with _ — {"location": {"city": "SF"}} becomes the facet location_city. An array field like "aliases": ["johndoe", "johnny"] is extracted as a single string facet.

A single log line can produce at most 50 facets by default; extraction fails silently beyond that limit. Raise it with skipAutoFacet:

logs-parser:
  kf_parsing_config:
    config: |-
      ...
      - skipAutoFacet:
        args:
          maxFacetsCount: 100
yaml

Once ingested, query the extracted facets from the Logs Explorer the same way as any other facet — see Logs Search Syntax.

Transform and fingerprinting

transform runs the exact same engine and action set as Relabel above — it’s the same function, just conventionally placed later in your config list, after the functions that extract the facets it promotes. The most common use is facet_to_label_map, promoting a facet extracted in Grammar (above) into a label — for example promoting a facet eventSource to a label source:

- transform:
    args:
      - action: "facet_to_label_map"
      - sourceLabels: "@eventSource"
      - targetLabel: "source"
    conditions:
      - matcher: "#source"
        op: "=="
        value: "awsLogSource"
yaml

Any other action from Relabel’s action table — drop, keep, label_map, and so on — works identically here; the name transform versus relabel reflects where you’ve placed it in the list, not a difference in capability.

Somewhere in the pipeline, the internal Kf-Parse stage also generates the log’s fingerprint — splitting the message into its static template and dynamic values, so that ts=2023-02-01T22:51:33Z caller=logging.go:29 method=Authorise result=false took=9.775µs becomes the fingerprint ts=<v_0> caller=<v_1> method=<v_2> result=<v_3> took=<v_4> plus the facets ts, caller, method, result, and took. Storing the template once and the variable values separately is what keeps Kloudfuse’s highly repetitive log traffic both fully indexed and storage-efficient — see Log Fingerprints for where this shows up in the Explorer. Kf-Parse itself isn’t configurable.

See also Log Timestamp Handling for how Kloudfuse selects and validates a log’s timestamp specifically.

Other pipeline functions

Beyond remap, relabel/transform, and dissect/grok, the pipeline supports these additional functions — each a list item under config, exactly like the examples above:

Function What it does

addFacet

Adds a facet with a fixed value. Args: name, value, and optional replace (boolean, default true — whether to overwrite an existing facet of the same name).

removeFacets

Removes one or more named facets. Args: names (a list).

renameFacet

Renames a facet. Args: source (existing facet name), target (new name).

moveLabelToFacet

Moves a label’s value into a facet of the same or a different name. Args: name (the label).

setLogLevel

Forces the log’s level to a fixed value, overriding whatever was detected. Args: level.

setTimestamp

Parses a facet as the log’s timestamp using one or more explicit formats, instead of relying on automatic timestamp detection. Args: facet, formats (a list).

configureTsExtract

Tunes how far into the message the automatic timestamp-detection heuristic looks. Args: lookaheadLength.

whitelistFacets

Restricts extraction to only the named facets, discarding everything else auto-detection would otherwise extract. Args: facets (a list).

logFmt

Parses the message as logfmt (space-separated key=value pairs) instead of relying on JSON auto-parsing or a custom grammar.

dropLogLine

Unconditionally drops the log line — pair it with conditions to drop only lines matching a condition, the same way you’d scope any other function.

archive

Controls this log’s archival behavior — which archive it’s written to, and whether it’s indexed for live search as well. See Archiving Logs and Hydration.

A further set of low-level tokenization primitives — resetPos, truncate, tokenize, rebuild, tokenizedKvExtract, tokenizedTsExtract, dissectKVOptimize, and tokenizedNewr — exist for advanced custom parsing scenarios beyond what dissect and grok cover directly. They’re implementation details rather than something most configurations need; contact your Kloudfuse representative if dissect or grok don’t cover a format you’re trying to parse.