Isolationforest Processor
contrib
Maintainers: @atoulme, @aarvee11
Source: opentelemetry-collector-contrib
Supported Telemetry
Overview
✨ Key Features
⚙️ How it Works
-
Training window – The processor keeps up to
window_sizeof the most recent data points for every feature‑group. -
Periodic (re‑)training – Every
training_interval, it drawssubsample_sizepoints from that window and growsforest_sizerandom isolation trees. - Scoring – Each new point is pushed through the forest. Shorter average path length ⇒ higher anomaly score.
- Adaptive sizing – When enabled, window size automatically adjusts based on traffic velocity, memory usage, and model stability.
-
Post‑processing –
- If
add_anomaly_score: true, a gauge metriciforest.anomaly_scoreis emitted with identical attributes/timestamp. - If the score ≥
anomaly_threshold, the original span/metric/log is flagged withiforest.is_anomaly=true. - If
drop_anomalous_data: true, flagged items are removed from the batch instead of being forwarded.
- If
Contamination rate – instead of hard‑codingPerformance is linear inanomaly_threshold, you can supplycontamination_rate(expected % of outliers). The processor then auto‑derives a dynamic threshold equal to the(1 – contamination_rate)quantile of recent scores.
forest_size and logarithmic in window_size; a default of 100 trees and a 1 k‑point window easily sustains 10–50 k points/s on a vCPU.
🔧 Configuration
🔄 Adaptive Window Configuration
When enabled, the processor automatically adjusts window size based on traffic patterns and resource constraints:
See the sample below for context.
📄 Sample config.yml
Note: Useroutingconnectorto seggregate the different kind of spans(db, messaging etc.) and send them to separateisolationforestprocessordeployments so the anomaly detection is pertianing to the respective category of signals.
What the example does
🚀 Best Practices
- Tune
forest_sizevs. latency – start with 100 trees; raise to 200–300 if scores look noisy. - Use per‑entity models – add
features(service, pod, host) to avoid global comparisons across very different series. - Let contamination drive threshold – set
contamination_rateto the % of traffic you’re comfortable labelling outlier; avoid hand‑tuninganomaly_threshold. - Use adaptive window sizing – enable for dynamic workloads; the processor will automatically grow windows during high traffic and shrink under memory pressure.
- Route anomalies – keep
drop_anomalous_data=falseand add a simple [routing‑processor] downstream to ship anomalies to a dedicated exporter or topic. - Monitor model health – the emitted
iforest.anomaly_scoremetric is perfect for a Grafana panel; watch its distribution and adapt window / contamination accordingly.
🏗️ Internals (High‑Level)
training_interval
Scoring cost: O(forest_size × log subsample_size) per item
Note: With adaptive window sizing enabled, current_window_size dynamically adjusts between min_window_size and max_window_size based on traffic patterns and memory constraints, making training costs adaptive to workload conditions.
🤝 Contributing
- Bugs / Questions – please open an issue in the fork first.
- Recently added: Adaptive window sizing for dynamic traffic patterns.
-
Planned enhancements
- Multivariate scoring (multiple numeric attributes per point).
- Expose Prometheus counters for training time / CPU cost.
Configuration
Example Configuration
Last generated: 2026-08-24