A blueprint for production RAG

Building RAG for large-scale, customer-facing applications requires balancing answer accuracy, retrieval depth and response time across complex, changing datasets. As AI agents make repeated retrieval calls within a task, these trade-offs become harder to manage.

Drawing on our experience supporting large-scale RAG applications, the RAG Blueprint turns practical lessons into a modular application template built on the Vespa AI Search Platform. It guides developers and architects through key decisions: how to structure searchable information, evaluate retrieval and ranking quality, and configure retrieval strategies for different tasks.

Practitioners can adapt the Blueprint to their own data and use cases, test their design choices and tune the application for production.

The RAG Blueprint overview

Make informed design decisions

Explore the choices that shape retrieval quality, response time and scalability. The RAG Blueprint explains the trade-offs and shows how to apply and evaluate each approach for your application.

Use the arrows to browse the six design decisions, with links to implementation guidance for each.

Benefits of The RAG Blueprint

  • Improve retrieval accuracy

    Apply proven retrieval and ranking practices, evaluate them against your own data and refine the context passed to the LLM for more relevant answers.

  • Accelerate production development

    Start with a modular application template that turns production experience into reusable code and design guidance, while preserving control over your implementation.

  • Build on a scaleble foundation

    Adapt an architecture designed to scale with your data and query load, with guidance for balancing retrieval quality, latency, and computational cost.

Perplexity uses Vespa to power fast, accurate, and trusted answers for millions of users

Perplexity’s search and answer experiences illustrate the demands placed on AI applications: enormous information collections, rapidly changing content, and users who expect fast, relevant answers. As these applications evolve, retrieving the right evidence remains essential to delivering useful, reliable responses.

Next steps

  • Tutorial

    Read The RAG Blueprint tutorial in the documentation,

    Read the documentation
  • Python notebook

    This notebook goes into more detail through The RAG Blueprint code. Learn how to set up machine learned ranking and how to conudtc rigourous testing and evaluation of a RAG application.

    Access the Python notebook
  • Sample application

    Follow the steps and deploy a trained, evaluated and fully functional RAG application.

    Go to GitHub
  • Getting started

    Or, if you are ew to Vespa and want to know how to get started and learn the basics.

    Get started

Frequently asked questions

Need more than a quick answer?

If these FAQs don't answer your question, there are several ways to continue:

Learn the fundamentals with our free online training at learn.vespa.ai.

Experience Vespa yourself with a free Vespa Cloud trial.

Watch the Getting Started with Vespa AI Search YouTube video

Contact our team to discuss your application or migration project.
Q: What is The RAG Blueprint?
The RAG Blueprint is a best-practice template based on real-world Vespa deployments, designed to help teams implement RAG systems efficiently and reliably. It provides practical guidance on retrieval and ranking, including support for hybrid search that combines keyword, vector, and structured signals. The Blueprint integrates with common embedding models and leverages Vespa’s built-in features, such as phased retrieval and query execution optimizations, to ensure low-latency performance at scale, even across billion-document workloads.
Q: What is the difference between RAG, The RAG Blueprint, and Retrieval Engineering?
RAG (retrieval-augmented generation) is an application pattern that retrieves relevant information and provides it to an LLM to help generate informed answers.
The RAG Blueprint is a practical reference application and tutorial that helps engineers explore how to build and evaluate RAG with Vespa, including decisions about data structure, retrieval, and ranking.
Retrieval Engineering is the broader discipline of designing, evaluating, and operating systems that retrieve and rank the right information. It addresses relevance, freshness, performance, and scale across RAG, AI agents, search, recommendations, and personalization.
The RAG Blueprint shows how Retrieval Engineering principles apply to a RAG application.
Q: What problem does it solve?
Deploying large-scale RAG systems in production presents several challenges that go beyond proof-of-concept implementations. At scale, maintaining accurate retrieval becomes harder as data volume and variety increase. Systems must combine keyword, semantic, and metadata-based signals to ensure relevant results, especially when content is noisy or domain-specific. Latency and throughput are also critical, as RAG pipelines must handle complex query chains and model inference in real-time, often across billions of documents.
These challenges become even more pronounced as research deepens, where LLMs must issue multiple queries, evaluate intermediate results, and reason across sources to produce trustworthy answers. This increases the demand on retrieval infrastructure, which must support high query rates, tight latency budgets, and rapid updates. Enterprises also face operational hurdles, such as enforcing access controls, keeping indexes up to date, and managing the costs of running embedding models and LLMs at scale. Without the right architecture, deep research use cases can expose the limits of traditional vector databases, delaying deployment and reducing the effectiveness of GenAI initiatives.
Q: Who is The RAG Blueprint for?
The RAG Blueprint is intended for engineers proficient in Vespa developing production-ready RAG applications. By following a predefined application, it provides a series of steps for developing a Vespa RAG application, including validating your system, demonstrating how to implement machine-learning document ranking, and outlining how to configure Vespa for optimal performance.
Q: Can I use The RAG Blueprint for my use case?
The RAG Blueprint provides a sample application and tutorial for learning and evaluating RAG approaches with Vespa. You can apply its design patterns and techniques to your own application, adapting data structures, retrieval methods, and ranking to your requirements.
The sample application is a learning resource rather than a ready-to-deploy solution for your specific use case.
Q: Is The RAG Blueprint for Vespa Cloud only?
No. The RAG Blueprint can be used with both Vespa Cloud and self-managed (open source) Vespa deployments. However, certain features, such as advanced chunking, require Vespa version 8.543.14 or later. Vespa Cloud handles upgrades automatically, while self-hosted users must manage them manually. Vespa Cloud also simplifies secure integration with off-the-shelf LLMs by letting you store API keys in the built-in secret store, avoiding the need to pass keys in request headers.
Q: What is included in the Blueprint package?
Sample Application
Tutorial Documentation
- How to structure an application
- How to specify ranking
- How to use ML to learn the best ranking formula
- How to use validation queries to assess retrieval performance
Overview Video
Python notebook