Create an APM alert
APM alerts watch a service metric — request rate, latency, error rate, or Apdex — and fire when it crosses a threshold you set, or deviates from its own learned baseline. Open the wizard from Alerts > New Alert, choosing APM as the alert type, or from the No APM alerts + control on a service’s Service detail page, which opens the wizard already scoped to that service. The wizard has four steps.
1. Choose a detection method
- Threshold
-
Fires when the metric crosses a fixed value you set. Use this for anything you can put a hard number on — for example, error rate should never exceed 1%, or an SLA-bound endpoint should never exceed 500ms p99.
- Anomaly
-
Fires when the metric deviates from its own learned baseline rather than a fixed number. Use this for services whose normal traffic and latency vary a lot by time of day or day of week, where a fixed threshold would either miss a real regression during a quiet period or false-positive during an expected peak.
2. Pick a service or write a query
- Service
-
Select the service to alert on.
- Alert on
-
The metric to watch: Requests/s, Errors/s, Error rate, p50/p75/p90/p95/p99 latency, Average latency, or APDEX. Prefer a percentile (p90 or p99) over an average for latency alerts — an average hides the slow tail that’s usually what users actually experience.
- Additional filters
-
Optionally narrow the alert to part of the service — one endpoint, or one span type (for example,
span_type=web) — instead of the whole service.
3. Set the alert condition
Both detection methods share the same base condition:
- When
-
The aggregation window to evaluate — for example, Last.
- of query
-
Which query the condition applies to (
A, or another letter if you’re alerting on a combined/multi-query expression). - Operator
-
The comparison: is above, is below, and so on.
- Threshold
-
The value to compare against. For a Threshold alert, this is the number the metric must cross. For an Anomaly alert, this is applied to the deviation the algorithm computes, not the raw metric.
Anomaly algorithm options
If you chose Anomaly in step 1, also configure:
- Algorithm
-
Basic (rolling quantile), Agile (SARIMA), Robust (seasonal decomposition), or Agile + Robust (Prophet). See each algorithm’s page for how it models a baseline and when it fits your traffic pattern.
- Bound
-
Sensitivity, from 1 (most sensitive) upward. A lower bound flags smaller deviations — and produces more alerts.
- Band
-
Which side of the baseline to watch: upper, lower, or both. Use upper for a metric where only increases are bad (latency, error rate); use both for a metric where either direction is meaningful (request rate dropping to zero is as much a problem as it spiking).
- Learn from
-
How much recent history the algorithm uses to establish its baseline.
- Repeats every
-
Shown for Robust and Agile + Robust, which model a recurring seasonal pattern (for example, Hourly) rather than only recent history — use this when a service’s normal traffic follows a daily or weekly cycle.
Optional settings
Whichever detection method you chose, tune how the alert behaves before you rely on it:
- Warning and recovery thresholds
-
A less severe threshold that notifies without paging, and the value the metric must return to before the alert is considered resolved.
- Evaluation frequency
-
How often Kloudfuse re-checks the condition.
- No data and error handling
-
What the alert does if the query returns no data or fails to evaluate — for example, whether missing data itself should be treated as a problem.
4. Add alert details
Name the alert for the problem it represents, not the query — it’s read by whoever gets paged, not by you. Choose the folder the alert rule is stored in (folders also carry permissions), then configure the notification routing (contact points) that should be notified when it fires.
Service-level objectives
Service-level objectives (SLOs) complement threshold and anomaly alerts for services where you want to track compliance against a target over a rolling window (for example, "99.9% of requests under 500ms over 30 days") rather than react to individual spikes. Configure one from the No SLO + control on the Service detail page; once configured, SLO burn appears in that page’s SLOs tab. Reserve SLOs for the handful of services with an actual reliability commitment — an SLO on every service just becomes noise alongside your threshold and anomaly alerts.