<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/how-hubspot-scaled-semantic-search-to-20-billion-vectors-vykkjoxfy" -->

---
title: How HubSpot Scaled Semantic Search to 20 Billion Vectors
description: HubSpot&#x27;s engineering team describes how their internal Vector as a Service (VaaS) platform scaled from a proof of concept to over 20 billion vectors across...
canonical: https://daily.dev/posts/how-hubspot-scaled-semantic-search-to-20-billion-vectors-vykkjoxfy
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: How HubSpot Scaled Semantic Search to 20 Billion Vectors | daily.dev
og:description: HubSpot&#x27;s engineering team describes how their internal Vector as a Service (VaaS) platform scaled from a proof of concept to over 20 billion vectors across...
og:url: https://daily.dev/posts/how-hubspot-scaled-semantic-search-to-20-billion-vectors-vykkjoxfy
og:image: https://api.daily.dev/og/posts/VYkkjOxfy.png
og:image:alt: How HubSpot Scaled Semantic Search to 20 Billion Vectors
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# How HubSpot Scaled Semantic Search to 20 Billion Vectors

**[InfoQ](https://daily.dev/sources/infoq)** · 4 min read · 1 upvotes · 0 comments

## Summary

HubSpot's engineering team describes how their internal Vector as a Service (VaaS) platform scaled from a proof of concept to over 20 billion vectors across 38+ teams. Built on top of Qdrant running on-premises, VaaS adds access control, embeddings generation, data versioning, and feedback collection. The team migrated from Helm-based manual cluster management to a custom Kubernetes Operator framework with 'Translators' that reconcile desired and actual cluster state every 60 seconds, automating shard movement, replication recovery, and cluster lifecycle. This reduced cluster spin-up time from hours to minutes and eliminated standby clusters. The platform now spans 200+ indexes, 140+ clusters, five regions, and handles write traffic peaking at 100,000 requests per second, supporting agents, RAG, and contact deduplication use cases.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://www.infoq.com/news/2026/07/hubspot-semantic-vector-search>

## Questions this post answers

### Why did HubSpot choose Qdrant for its vector search infrastructure at scale?

HubSpot chose Qdrant because it can run on-premises, supports named vectors, hybrid search, multi-stage querying, and weighted reranking, and offers cost controls like quantisation and on-disk storage. Running it in-house also let HubSpot integrate internal tracing, cost tracking, rate limiting, scaling, and security tooling while retaining control over customer data.

_daily.dev surfaces engineering deep dives like this for teams weighing vector database options._

### How did HubSpot automate management of its Qdrant clusters at massive scale?

HubSpot moved from manual Helm-based deployments to an internal Kubernetes Operator framework built around components called Translators, which reconcile desired and actual cluster state every 60 seconds. This automated cluster creation, decommissioning, shard movement, and replication recovery, cutting cluster spin-up time from hours to minutes and removing the need for standby clusters.

_Engineers automating infrastructure at scale can follow similar operator patterns via daily.dev._

### How many vectors and clusters does HubSpot's vector search platform manage?

HubSpot's Vector as a Service platform manages more than 20 billion vectors across 38-plus teams, spanning over 200 indexes, 140-plus clusters, five regions, and two environments. Write traffic peaks at 100,000 requests per second, and the system supports agents, retrieval-augmented generation, and contact deduplication use cases.

_Teams benchmarking retrieval infrastructure at scale can track real numbers like these on daily.dev._

## Similar posts on daily.dev

- [The story of BigQuery vector search](https://daily.dev/posts/the-story-of-bigquery-vector-search-tmmmlqwoo) · Google Cloud · 1 upvotes · 0 comments
- [How Data 360 Vector Search Delivers Near Real-Time Intelligence](https://daily.dev/posts/how-data-360-vector-search-delivers-near-real-time-intelligence-cm0fjdii6) · Salesforce Engineering · 2 upvotes · 0 comments
- [Amazon S3 Vectors now generally available with increased scale and performance](https://daily.dev/posts/amazon-s3-vectors-now-generally-available-with-increased-scale-and-performance-0gpqkus8w) · AWS · 0 upvotes · 0 comments

---

Tags: [#kubernetes](https://daily.dev/tags/kubernetes), [#rag](https://daily.dev/tags/rag), [#vector-search](https://daily.dev/tags/vector-search)

[View this post on daily.dev](https://daily.dev/posts/how-hubspot-scaled-semantic-search-to-20-billion-vectors-vykkjoxfy)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"How HubSpot Scaled Semantic Search to 20 Billion Vectors","url":"https://daily.dev/posts/how-hubspot-scaled-semantic-search-to-20-billion-vectors-vykkjoxfy","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/how-hubspot-scaled-semantic-search-to-20-billion-vectors-vykkjoxfy"},"datePublished":"2026-07-07T08:05:47.581Z","dateModified":"2026-07-07T10:21:43.125Z","description":"HubSpot's engineering team describes how their internal Vector as a Service (VaaS) platform scaled from a proof of concept to over 20 billion vectors across...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/53c78f5c89dd27aa98b77d41447665d9?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/53c78f5c89dd27aa98b77d41447665d9?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"InfoQ","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"InfoQ","logo":"https://media.daily.dev/image/upload/t_logo,f_auto/v1/logos/afc3bced3e1e4b188dd9127017a60e0c","url":"https://daily.dev/sources/infoq"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/how-hubspot-scaled-semantic-search-to-20-billion-vectors-vykkjoxfy","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":1},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"kubernetes,rag,vector-search","timeRequired":"PT4M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"InfoQ","item":"https://daily.dev/sources/infoq"},{"@type":"ListItem","position":3,"name":"How HubSpot Scaled Semantic Search to 20 Billion Vectors"}]}
```

