<200ms
Neural search latency with Exa Instant, reduced
As published on zilliz.com. Captured by usedby on Oct 6, 2026.
What happened
Exa uses Zilliz Cloud (managed Milvus) to power its entity search layer for companies, people, and code, serving as the primary index and recency cache. It handles hybrid dense and sparse vector search with metadata filtering and frequent upserts, feeding Exa's Search API and Websets product.
Summary written by usedby from the source page, in English. The figures are those of Zilliz and Exa, not ours.
Hybrid search combining dense vectors, sparse vectors, RRF reranking, and metadata filters in a single API call. Exa Instant reduced neural search latency from seconds to under 200ms
Zilliz Cloud has been an important part of Exa’s journey to build and scale entity search, giving us the retrieval performance and operational simplicity we need to scale quickly and confidently.
Zilliz gave us real-time retrieval for our AI search system at scale with tight latency targets. It freed up engineering cycles and let us focus on improving reasoning on the model side, not managing infrastructure.
We believe AI agents will become a fundamental interface for how people work, learn, and make decisions, and that only happens if those systems can access real-world information with speed, precision, and trust.
What the story claims, and what we checked
We compared the story with its live page on Oct 6, 2026.
- The figure: <200msCheckedThe figure is on the page; its label is our wording.
- The passage quoted aboveCheckedCopied word for word from the page, near the name of Exa.
- Exa uses ZillizCheckedConfirmed line. Latest check across sources: Oct 6, 2026.
- The result itselfNot checkedWe quote it; we did not measure it.




