1.4 Retrieval systems
Today’s foundation models are trained on enormous mixtures of web pages, books, code, papers, and other data. That breadth lets them answer many general questions by pattern completion, but it is neither the entire written record nor a complete, current knowledge base. What is built in is still not the knowledge an organization needs at run time: this company’s contracts, this codebase’s architecture decisions, yesterday’s incident review, what this customer was promised on Tuesday. That is the gap retrieval systems exist to fill — supplying, at run time, the material the model was never trained on.
The retrieval layer decides what an inference engine can be shown during a run, and in 2026 it is a mature category with a distinct shape: embeddings turn text into vectors, a database indexes them, and a query finds the nearest neighbors and hands them to the model. Retrieval answers “what should this run be able to see?” Memory, which follows, answers “what should this system remember?” They are related — many memory systems use retrieval infrastructure underneath — but the categories are not the same, and conflating them is how systems end up with a vector database doing a memory system’s job.
1.4.1 The vector database layer
A vector database stores learned numerical representations. Text — a paragraph, a document, a support ticket — is passed through an embedding model that converts it into a long list of numbers, a vector, arranged so that texts with similar meanings land near each other in that number space. The database indexes the vectors and answers one question very fast: of everything stored, which items are nearest to this query? That is semantic search — finding matches by similarity rather than by keyword. The embedding is not a literal encoding of meaning, and it does not store the source text or everything known about it.
An orchestrator may never operate a vector database directly. They usually sit one layer down: inside a memory system (Section 5.5), inside a retrieval pipeline a framework assembled, or behind an agent’s search tool. Understanding the layer still matters, because what it can do shapes what the layers above can promise — whether memory can scale, whether search can filter by permission, whether the index survives constant writes.
The category became visible almost overnight. When ChatGPT arrived, developers needed a way to ground model answers in their own documents, and the pattern that emerged was retrieval-augmented generation — RAG: retrieve the relevant passages, put them into the model’s context, and generate the answer from those passages instead of from the model’s memory. Vector databases were the retrieval engine underneath, and the market reacted immediately. Pinecone raised one hundred million dollars at a $750 million valuation in 2023, its CEO describing demand that took off overnight; searches for “vector database” grew roughly eleven-fold between early 2023 and early 2025; Milvus, Chroma, Weaviate, and Qdrant all became visible names within a year.21
The current field:
| System | What it is | License | What distinguishes it |
|---|---|---|---|
| Pinecone | Managed serverless vector database | Proprietary | No-operations deployment; in 2026 it has climbed the stack with managed RAG and a “knowledge engine for agents” (Nexus) |
| Weaviate | Open-source vector database | BSD 3-Clause | Hybrid search (keyword + vector) as a first-class feature; an open-source core with a managed cloud |
| Qdrant | Open-source vector database in Rust | Apache 2.0 | Performance and resource efficiency; strong filtering; a 2026 roadmap aimed at agent-native retrieval |
| Milvus | Open-source distributed vector database | Apache 2.0 | Built for billion-scale; a Linux Foundation AI & Data project, with managed Zilliz Cloud |
| Chroma | Open-source embedding database | Apache 2.0 | Developer-friendliness first; the fastest path from zero to a working prototype, with a managed cloud since 2025 |
| pgvector | PostgreSQL extension | PostgreSQL | Vector search as a feature of the database you already run — inherit Postgres backups, permissions, and transactions |
Two readings of that table matter. The first is the one most teams act on: pick the vector database by scale and operations model, from a laptop prototype on Chroma to a billion-vector estate on Milvus. The second is the trend line: vector search is becoming a feature of mainstream databases rather than a standalone category. pgvector puts it inside PostgreSQL; DynamoDB added native vector search in August 2026; the major platforms now ship some form of it.
That is what makes this a manageable concern rather than a strategic one. An orchestrator choosing retrieval infrastructure in 2026 is not betting a platform for the next decade; the category is commoditizing underneath whatever choice is made. The durable questions are the ones this chapter keeps asking of every layer: who owns the index, what goes into it, when was it last refreshed, and what does it let each run see.
1.4.2 The retrieval tooling
Around the databases sits the tooling that assembles retrieval into working pipelines. This is the layer most orchestrators actually encounter, because the frameworks here make the decisions a raw vector database leaves open — how documents are parsed and chunked, how they are embedded, how queries are rewritten, how results are ranked — and package the answers as configurable components:23
- LlamaIndex — the data-ingestion specialist: deep document parsing (including tables and scanned files), many indexing strategies, and integrations with more than forty vector stores. The usual choice when the documents are messy.
- LangChain and LangGraph — the composability generalists: loaders, splitters, embedders, retrievers, and rerankers as building blocks that chain into pipelines or graph workflows. Best when retrieval is one part of a larger orchestrated flow.
- Haystack 3.0 (deepset) — production pipelines with evaluation built in; its July 2026 release puts agents at the center while keeping typed, testable components. The choice of teams that need to prove retrieval quality, not just ship it.
- DSPy — a different idea: instead of hand-writing prompts and retrieval settings, declare the pipeline and optimize it against a metric. It treats prompting as a programming problem.
- RAGAS — evaluation rather than construction: measures whether a retrieval system actually works — faithfulness, answer relevance, context precision — so the rest of this list can be improved honestly.
- The hosted services — Pinecone Assistant, the model vendors’ file-search and web-search tools, and managed RAG platforms — fold chunking, embedding, and query planning into one API call. Fastest path to working, least control over how.
For learning the area, start with the retrieval documentation of whichever framework you adopt — LlamaIndex and Haystack both maintain full retrieval guides — and read the RAGAS documentation to learn how retrieval quality is measured. The durable skill is knowing which decisions exist (chunking, embedding model, ranking, filtering) and asking of any tool who owns each one.
There is a straightforward way to say what retrieval means for an orchestrator: what a system can retrieve, especially from memory, is often also a permissions question. Which documents an agent may search, which repositories it may index, which customer records it may pull — those decisions are already permission decisions, and they belong in the same review as the rest of the access model.
1.4.3 Where retrieval shows up for an orchestrator
Retrieval shows up in three places in an orchestrated system. The first is context assembly: the orchestrator decides what a run can see — which documents, which repositories, which customer records — and retrieval is the machinery that executes that decision. The second is inside memory: many memory products run on retrieval infrastructure underneath, so the quality of the memory system is partly the quality of its retrieval — while others also use files, relational or key-value stores, graphs, or provider-managed state. The third is as agent tools: a search tool, a documentation lookup, a web research capability — each is retrieval wrapped in a permission grant.
An orchestrator tuning chunk size or the embedding model is making a quality decision that shows up in every answer the system gives. An orchestrator reading a trace — what did the system actually retrieve before it answered? — is doing the forensic half of the job.
Vector-database boom: Pinecone’s $100M Series B at a $750M valuation (April 2023, led by Andreessen Horowitz) per VentureBeat, “AI startup Pinecone raises $100 million as vector database market for LLMs heats up,” https://venturebeat.com/ai/ai-startup-pinecone-raises-100-million-as-vector-database-market-for-llms-heats-up, which quotes CEO Edo Liberty on demand that “took off overnight” after ChatGPT; the concurrent Chroma/Weaviate/Pinecone term-sheet wave per Business Insider, “Chroma, Weaviate, and Pinecone Raise Funding From A16z, Index Ventures,” March 2023, https://www.businessinsider.com/chroma-weaviate-pinecone-raise-funding-a16z-index-vector-database-ai-2023-3. The ~11x growth in “vector database” searches between January 2023 and January 2025 per DataAspirant, “Most Popular Vector Databases You Must Know,” https://dataaspirant.com/popular-vector-databases. All figures are as reported by the cited outlets.↩︎
Vector systems as of August 29, 2026: Pinecone, https://www.pinecone.io; Weaviate, https://weaviate.io; Qdrant, https://qdrant.tech; Milvus, https://milvus.io; Chroma, https://www.trychroma.com; and pgvector, https://github.com/pgvector/pgvector. General databases are also adding vector search; DynamoDB announced it in August 2026, https://aws.amazon.com/about-aws/whats-new/2026/08/amazon-dynamodb-vector-search/. Product features and licenses vary.↩︎
Retrieval-tooling landscape as of August 2026: LlamaIndex (https://www.llamaindex.ai) — document parsing, indexing strategies, 40+ vector-store integrations; LangChain/LangGraph retrieval components (https://docs.langchain.com); Haystack 3.0 (deepset, Apache 2.0; released July 20, 2026, agents-first with pipelines and evaluation retained) per https://www.turingpost.com/p/rag-tools and https://haystack.deepset.ai; DSPy (https://dspy.ai) — programmatic pipeline optimization; RAGAS (https://docs.ragas.io) — retrieval evaluation metrics (faithfulness, answer relevance, context precision/recall); hosted RAG services (Pinecone Assistant, https://www.pinecone.io). Category surveys used for placement: Firecrawl, “15 Best Open-Source RAG Frameworks in 2026,” https://www.firecrawl.dev/blog/best-open-source-rag-frameworks; AY Automate, “10 Best RAG Frameworks and Libraries in 2026,” https://www.ayautomate.com/blog/best-rag-frameworks; Meilisearch, “10 best RAG tools and platforms,” https://www.meilisearch.com/blog/rag-tools. Placement characterizations are this book’s synthesis of those surveys, not vendor claims.↩︎
- 1.1 The technology under discussion
- 1.2 The survey: connecting the taxonomy to the market
- 1.3 Inference engines and models
- 1.4 Retrieval systems
- 1.5 Memory systems
- 1.6 Tool interfaces and actuation
- 1.7 Identity and access control
- 1.8 Policy and permission systems
- 1.9 Durable orchestration engines
- 1.10 Task agents
- 1.11 Everything is called an agent
- 1.12 Multi-agent coding systems
- 1.13 The architecture of participation
- 1.14 The chapter in one picture