Most RAG projects should start with a search and a cited answer. Add hybrid search when exact terms matter, a graph when relationships matter, and an agent when the next search depends on the previous result. Here is how to choose.
Retrieval-augmented generation (RAG) gives a model external evidence before it answers. The patterns below change how that evidence is found or checked. They can be combined: an agent can use hybrid search, retrieve a diagram, and check its sources.
Which one would I use?
- Basic RAG: “How do I rotate a service token?” Retrieve the maintained runbook and cite it. Start here when a few passages contain the answer. Fix parsing and chunking before adding complexity.
- Hybrid RAG: “What does ERR-742 mean in version 3.2?” Combine keyword and vector search, filter by version, then rerank. Exact identifiers need more than semantic similarity. Reranking adds latency and cannot rescue evidence the search never found.
- HyDE: “Why does login disappear after some idle time?” Generate a hypothetical passage about session expiry and use it to find real sources. Try simpler query rewriting first. The invented draft guides retrieval; it is never evidence. HyDE paper.
- GraphRAG: “Which products depend on this supplier?” Follow supplier → component → product relationships. Useful for dependencies and lineage; less useful for simple lookups. Extracted relationships need validation, especially when completeness matters.
- Multimodal RAG: “Which valve is upstream of this pump?” Retrieve the schematic, not just its OCR text. Use it when layout, images, audio, or video carry the answer. Preserve page or timestamp references.
- Corrective RAG: “Which procedure applies? These versions disagree.” Check the evidence, retrieve the current revision, or surface the conflict. Bound retries. A model approving its own answer is not proof. CRAG paper.
- Adaptive RAG: One assistant handles greetings, lookups, and investigations. Route each request to the appropriate path and budget. Test the router: skipping retrieval must not become permission to guess company facts.
- Agentic RAG: “Why did this incident recur after last week's fix?” Retrieve the incident, identify the fix, inspect changes, then search again if needed. Use it for investigations whose steps depend on discoveries. Cap time and tool calls.
A naming note: advanced RAG groups improvements such as rewriting and reranking; modular RAG describes swappable components. Self-RAG is a specific trained reflection approach, not simply a “double-check your answer” prompt. These labels describe different aspects of a system, not competing product tiers.
What can AWS do today?
Checked 5 October 2026. Choose between customer-managed Knowledge Bases, where you control more of the retrieval stack, and Bedrock Managed Knowledge Bases, launched in June 2026 with managed ingestion, storage, hybrid search, ranking, and agentic retrieval. Less infrastructure to operate means fewer tuning choices.
| You need | AWS option and important limit |
|---|---|
| Document answers | RetrieveAndGenerate, or Retrieve plus your own answer stage. |
| Hybrid + reranking | For customer-managed KBs, HYBRID needs a supported RDS, OpenSearch Serverless, or MongoDB store with a filterable text field. Bedrock reranking is text-only. |
| GraphRAG | Neptune Analytics integration: S3 sources, no graph-build customization or graph autoscaling. |
| Images, audio, video | Multimodal Knowledge Bases. Native embeddings and text conversion have different model, retrieval, and generation compatibility. |
| Iterative research | Managed KB AgenticRetrieveStream plans, retrieves, and checks evidence. Custom HyDE, routing, and correction policies remain application work. |
Two easy mistakes: direct S3 Vectors integration provides semantic search, not hybrid search. And “How many customers are affected?” needs an authoritative database query, not a count of retrieved chunks. AWS offers structured querying for that. Check regional and model availability for your chosen combination.
Where RafiHive fits
RafiHive, our regulatory-affairs assistant, uses advanced RAG with rule-based evidence checks. Its implementation retrieves approved sources through Bedrock, can follow explicit regulatory cross-references with additional searches, and passes the evidence to the answer model. Retrieval scores help decide whether to answer, try an enabled web fallback, or report insufficient evidence; they do not prove correctness.
For example, a retrieved guidance passage may cite another regulation. RafiHive can fetch that referenced text before answering. This adds useful context without a graph database or an open-ended agent loop. The design priority is traceable evidence for an RA professional to review.
The trend, and my bet for what comes next
The visible AWS trend is managed agentic retrieval: planning and repeated searches are becoming service capabilities. September's document-permission debugging update also points toward the operational work behind reliable retrieval: understanding why a user can or cannot find a source.
My prediction for 2027–2028: hybrid search with reranking will be a common default, with bounded agentic retrieval for difficult questions. That balances answer quality, latency, and cost. Multimodal retrieval will grow in media-heavy products; GraphRAG will remain useful for relationship-heavy domains. This is an engineering forecast, not a measured popularity ranking.
Start simple. Add a retrieval step only when real questions show why you need it.
Compare candidates on the same questions: did they find the necessary evidence, cite it correctly, respect access permissions, and admit missing information? Then compare latency and cost per resolved question. A more elaborate diagram only earns its place if the answers improve.