Most RAG projects should start with a search and a cited answer. Add hybrid search when exact terms matter, a graph when relationships matter, and an agent when the next search depends on the previous result. Here is how to choose.

Retrieval-augmented generation (RAG) gives a model external evidence before it answers. The patterns below change how that evidence is found or checked. They can be combined: an agent can use hybrid search, retrieve a diagram, and check its sources.

Eight RAG patterns at a glance · open the image to zoom
Eight illustrated RAG flows: Basic searches once; Hybrid combines keyword and vector search with reranking; HyDE searches using a hypothetical draft; GraphRAG follows entity relationships; Multimodal retrieves media; Corrective retries weak retrieval; Adaptive routes different questions; Agentic plans and searches iteratively. Examples and selection guidance follow below.

Which one would I use?

A naming note: advanced RAG groups improvements such as rewriting and reranking; modular RAG describes swappable components. Self-RAG is a specific trained reflection approach, not simply a “double-check your answer” prompt. These labels describe different aspects of a system, not competing product tiers.

What can AWS do today?

Checked 5 October 2026. Choose between customer-managed Knowledge Bases, where you control more of the retrieval stack, and Bedrock Managed Knowledge Bases, launched in June 2026 with managed ingestion, storage, hybrid search, ranking, and agentic retrieval. Less infrastructure to operate means fewer tuning choices.

You needAWS option and important limit
Document answersRetrieveAndGenerate, or Retrieve plus your own answer stage.
Hybrid + rerankingFor customer-managed KBs, HYBRID needs a supported RDS, OpenSearch Serverless, or MongoDB store with a filterable text field. Bedrock reranking is text-only.
GraphRAGNeptune Analytics integration: S3 sources, no graph-build customization or graph autoscaling.
Images, audio, videoMultimodal Knowledge Bases. Native embeddings and text conversion have different model, retrieval, and generation compatibility.
Iterative researchManaged KB AgenticRetrieveStream plans, retrieves, and checks evidence. Custom HyDE, routing, and correction policies remain application work.

Two easy mistakes: direct S3 Vectors integration provides semantic search, not hybrid search. And “How many customers are affected?” needs an authoritative database query, not a count of retrieved chunks. AWS offers structured querying for that. Check regional and model availability for your chosen combination.

Where RafiHive fits

RafiHive, our regulatory-affairs assistant, uses advanced RAG with rule-based evidence checks. Its implementation retrieves approved sources through Bedrock, can follow explicit regulatory cross-references with additional searches, and passes the evidence to the answer model. Retrieval scores help decide whether to answer, try an enabled web fallback, or report insufficient evidence; they do not prove correctness.

For example, a retrieved guidance passage may cite another regulation. RafiHive can fetch that referenced text before answering. This adds useful context without a graph database or an open-ended agent loop. The design priority is traceable evidence for an RA professional to review.

The trend, and my bet for what comes next

The visible AWS trend is managed agentic retrieval: planning and repeated searches are becoming service capabilities. September's document-permission debugging update also points toward the operational work behind reliable retrieval: understanding why a user can or cannot find a source.

My prediction for 2027–2028: hybrid search with reranking will be a common default, with bounded agentic retrieval for difficult questions. That balances answer quality, latency, and cost. Multimodal retrieval will grow in media-heavy products; GraphRAG will remain useful for relationship-heavy domains. This is an engineering forecast, not a measured popularity ranking.

Start simple. Add a retrieval step only when real questions show why you need it.

Compare candidates on the same questions: did they find the necessary evidence, cite it correctly, respect access permissions, and admit missing information? Then compare latency and cost per resolved question. A more elaborate diagram only earns its place if the answers improve.