<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/when-kpis-go-weird-anomaly-detection-with-python---juliana-ferreira-alves-rcfyr1qtq" -->

---
title: When KPIs Go Weird: Anomaly Detection with Python -...
description: A transcript of a PyCon talk introducing anomaly detection concepts and practical applications, covering why outlier detection matters (fraud prevention,...
canonical: https://daily.dev/posts/when-kpis-go-weird-anomaly-detection-with-python---juliana-ferreira-alves-rcfyr1qtq
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: When KPIs Go Weird: Anomaly Detection with Python - Juliana Ferreira Alves | daily.dev
og:description: A transcript of a PyCon talk introducing anomaly detection concepts and practical applications, covering why outlier detection matters (fraud prevention,...
og:url: https://daily.dev/posts/when-kpis-go-weird-anomaly-detection-with-python---juliana-ferreira-alves-rcfyr1qtq
og:image: https://api.daily.dev/og/posts/RcFyr1QTQ.png
og:image:alt: When KPIs Go Weird: Anomaly Detection with Python - Juliana Ferreira Alves
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# When KPIs Go Weird: Anomaly Detection with Python - Juliana Ferreira Alves

**[PyCon US](https://daily.dev/sources/pyconus)** · 24 min read · 0 upvotes · 0 comments

## Summary

A transcript of a PyCon talk introducing anomaly detection concepts and practical applications, covering why outlier detection matters (fraud prevention, healthcare, website bot traffic) and explaining unsupervised learning algorithms like Isolation Forest, One-Class SVM, and Local Outlier Factor. The speaker walks through a simplified scikit-learn example using fake height/weight/age data, discusses tuning the contamination parameter, and highlights real-world challenges like label scarcity, class imbalance, and the need to validate anomalies with business context before deploying such models.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://www.youtube.com/watch?v=7Zj9Ma2l0lM>

## Questions this post answers

### How does the isolation forest algorithm detect anomalies in data?

Isolation forest builds many decision trees that each try to isolate individual data points, then the trees vote on how likely each point is to be an outlier. Points that get separated from the rest of the data more quickly and easily, meaning they sit higher up in the tree, receive a higher anomaly score and are more likely to be flagged as outliers.

_daily.dev surfaces practical explainers like this for developers choosing an anomaly detection algorithm._

### What is the difference between isolation forest and local outlier factor for anomaly detection?

Isolation forest tries to find features that make a data point easy to separate from the rest, without profiling the whole data dimension, making it computationally lighter. Local outlier factor instead looks at the full set of dimensions (for example weight, height, and age together) and profiles similarity between points based on density, which makes it more computationally expensive but potentially more precise for subtle, non-obvious outliers.

_Developers weighing anomaly detection models can track trade-offs like these on daily.dev._

### Why is choosing the contamination parameter the hardest part of building an anomaly detection model with scikit-learn?

The contamination parameter requires estimating what percentage of your data is actually outliers, but in real-world data this percentage is unknown upfront. Practitioners typically test several contamination values and visually or experientially judge which produces results that make sense for the specific business problem, making model fine-tuning the most time-consuming part of the workflow.

_daily.dev helps practitioners compare real experiences fine-tuning models like these before shipping._

## Similar posts on daily.dev

- [A Practical Toolkit for Time Series Anomaly Detection, Using Python](https://daily.dev/posts/a-practical-toolkit-for-time-series-anomaly-detection-using-python-pzhoovjro) · Towards Data Science · 2 upvotes · 0 comments
- [Anomaly Detection on Industrial Sensor Data: Beyond Threshold Alerts](https://daily.dev/posts/anomaly-detection-on-industrial-sensor-data-beyond-threshold-alerts-13bgule7f) · CrateDB · 0 upvotes · 0 comments
- [5 Best Practices for Automated Anomaly Detection in Data Pipelines](https://daily.dev/posts/5-best-practices-for-automated-anomaly-detection-in-data-pipelines-2dx0liztg) · Decube · 0 upvotes · 0 comments
- [Anomaly detection on financial data — A classical ML based approach](https://daily.dev/posts/anomaly-detection-on-financial-data-a-classical-ml-based-approach-vnptvnil3) · Medium · 0 upvotes · 0 comments

---

Tags: [#machine-learning](https://daily.dev/tags/machine-learning), [#python](https://daily.dev/tags/python), [#scikit](https://daily.dev/tags/scikit), [#unsupervised-learning](https://daily.dev/tags/unsupervised-learning)

[View this post on daily.dev](https://daily.dev/posts/when-kpis-go-weird-anomaly-detection-with-python---juliana-ferreira-alves-rcfyr1qtq)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"When KPIs Go Weird: Anomaly Detection with Python - Juliana Ferreira Alves","url":"https://daily.dev/posts/when-kpis-go-weird-anomaly-detection-with-python---juliana-ferreira-alves-rcfyr1qtq","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/when-kpis-go-weird-anomaly-detection-with-python---juliana-ferreira-alves-rcfyr1qtq"},"datePublished":"2026-10-07T21:10:03.299Z","dateModified":"2026-10-07T21:10:24.735Z","description":"A transcript of a PyCon talk introducing anomaly detection concepts and practical applications, covering why outlier detection matters (fraud prevention,...","image":"https://i.ytimg.com/vi/7Zj9Ma2l0lM/sddefault.jpg","thumbnailUrl":"https://i.ytimg.com/vi/7Zj9Ma2l0lM/sddefault.jpg","isAccessibleForFree":true,"articleSection":"PyCon US","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"PyCon US","logo":"https://media.daily.dev/image/upload/s--DLedNxG1--/f_auto,q_auto/v1777800767/logos/pyconus?_a=BAMAMiWQ0","url":"https://daily.dev/sources/pyconus"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/when-kpis-go-weird-anomaly-detection-with-python---juliana-ferreira-alves-rcfyr1qtq","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":0},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"machine-learning,python,scikit,unsupervised-learning","timeRequired":"PT24M","video":{"@type":"VideoObject","name":"When KPIs Go Weird: Anomaly Detection with Python - Juliana Ferreira Alves","description":"A transcript of a PyCon talk introducing anomaly detection concepts and practical applications, covering why outlier detection matters (fraud prevention,...","thumbnailUrl":"https://i.ytimg.com/vi/7Zj9Ma2l0lM/sddefault.jpg","uploadDate":"2026-10-07T21:10:03.299Z","duration":"PT24M","url":"https://api.daily.dev/r/RcFyr1QTQ","embedUrl":"https://www.youtube.com/embed/7Zj9Ma2l0lM"}}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"PyCon US","item":"https://daily.dev/sources/pyconus"},{"@type":"ListItem","position":3,"name":"When KPIs Go Weird: Anomaly Detection with Python - Juliana Ferreira Alves"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/when-kpis-go-weird-anomaly-detection-with-python---juliana-ferreira-alves-rcfyr1qtq#faq","mainEntity":[{"@type":"Question","name":"How does the isolation forest algorithm detect anomalies in data?","acceptedAnswer":{"@type":"Answer","text":"Isolation forest builds many decision trees that each try to isolate individual data points, then the trees vote on how likely each point is to be an outlier. Points that get separated from the rest of the data more quickly and easily, meaning they sit higher up in the tree, receive a higher anomaly score and are more likely to be flagged as outliers. daily.dev surfaces practical explainers like this for developers choosing an anomaly detection algorithm."}},{"@type":"Question","name":"What is the difference between isolation forest and local outlier factor for anomaly detection?","acceptedAnswer":{"@type":"Answer","text":"Isolation forest tries to find features that make a data point easy to separate from the rest, without profiling the whole data dimension, making it computationally lighter. Local outlier factor instead looks at the full set of dimensions (for example weight, height, and age together) and profiles similarity between points based on density, which makes it more computationally expensive but potentially more precise for subtle, non-obvious outliers. Developers weighing anomaly detection models can track trade-offs like these on daily.dev."}},{"@type":"Question","name":"Why is choosing the contamination parameter the hardest part of building an anomaly detection model with scikit-learn?","acceptedAnswer":{"@type":"Answer","text":"The contamination parameter requires estimating what percentage of your data is actually outliers, but in real-world data this percentage is unknown upfront. Practitioners typically test several contamination values and visually or experientially judge which produces results that make sense for the specific business problem, making model fine-tuning the most time-consuming part of the workflow. daily.dev helps practitioners compare real experiences fine-tuning models like these before shipping."}}]}
```

