NVIDIA-AI-Blueprints/rag

This NVIDIA RAG blueprint serves as a reference solution for a foundational Retrieval Augmented Generation (RAG) pipeline.

What it solves

This project provides a reference architecture for building enterprise-grade Retrieval-Augmented Generation (RAG) pipelines. It addresses the need for AI responses that are grounded in trusted, real-time enterprise data to reduce hallucinations, ensure accuracy, and maintain compliance and freshness of information.

How it works

The system uses a modular orchestration layer (based on LangChain) to coordinate several GPU-accelerated components:

  • Data Ingestion: Extracts multimodal content (text, tables, charts, infographics, and audio) from documents using specialized NIM microservices.
  • Retrieval: Converts queries into embeddings and searches a GPU-accelerated vector database (such as Elasticsearch or Milvus) using hybrid search (dense and sparse).
  • Refinement: Employs reranking models to improve the precision of retrieved passages.
  • Generation: Uses LLMs (like Nemotron) to generate grounded responses, with optional reflection steps and programmable guardrails for safety.
  • Agentic RAG: For complex queries, it implements a LangGraph plan-and-execute pipeline that handles multi-hop and cross-document searches through scope discovery and parallel task execution.

Who it’s for

Developers and enterprises looking to deploy scalable, low-latency RAG systems that can handle multimodal data and complex, multi-step reasoning queries.

Highlights

  • Multimodal Support: Capable of extracting and processing data from tables, charts, and audio within documents.
  • Agentic Capabilities: Includes a plan-and-execute pipeline for ambiguous or multi-hop queries.
  • GPU Acceleration: Leverages NVIDIA NIM microservices and cuVS for high-performance indexing and search.
  • Enterprise Ready: Includes a reference UI, OpenAI-compatible APIs, and support for deployment via Docker or Kubernetes/OpenShift.
  • Evaluation Tools: Integrated RAGAS framework scripts for quality benchmarking.