TryHackMe Walkthrough: RAG Security Fundamentals
RAG Security Fundamentals introduces the security risks associated with Retrieval-Augmented Generation (RAG) systems.
RAG systems retrieve external information and provide it to an LLM as context before generating a response. While this improves the model's ability to work with up-to-date or private information, it also creates new attack surfaces.
A key security concern is inference-time data poisoning, where malicious or misleading documents influence the model's response without requiring the model itself to be retrained.
Other important risks include malicious content entering during ingestion, poisoned documents being selected during retrieval, and retrieved content manipulating the LLM through context injection.
Task 2 - RAG Architecture Overview
Overview:
This task explains the basic architecture of a RAG system and how data moves through it.The main components are:
Embedding Model: Converts text into numerical vectors that represent its meaning.
Vector Store: Stores document embeddings and enables similarity searches.
Retriever: Finds the most relevant documents based on the user's query.
LLM: Uses the retrieved documents as context to generate the final response.
The main security risks are concentrated around ingestion, retrieval, and context injection.
Questions & Answers
Embeddings
Embeddings represent the meaning of text as numerical vectors, allowing the system to perform similarity comparisons.
Retriever
The retriever searches the vector store and selects relevant documents based on similarity to the user's query.
Task 3 - RAG-Specific Attack Surface
Overview
RAG systems introduce unique attack surfaces because external documents directly influence the LLM during inference.
The main areas are:
Document ingestion: Malicious or untrusted documents can enter the knowledge base.
Embedding generation: Documents are converted into numerical vectors, making security context harder to inspect.
Similarity-based retrieval: Documents are selected based on semantic relevance rather than trustworthiness.
Context injection: Retrieved content is passed directly into the LLM's context.
Questions & Answers
1. Which RAG stage introduces the largest indirect attack surface?
Similarity-based retrieval2. What component is lost during embedding generation that affects security?
Context
Task 4 - Retrieval Abuse & Context Manipulation
Overview
Retrieval abuse is when an attacker influences which documents a RAG system retrieves, causing malicious or misleading content to enter the model’s context. This can happen through:- Passive poisoning: Malicious content is added to the knowledge base and waits to be retrieved.
- Active manipulation: Content is deliberately crafted to rank highly for specific/sensitive queries.
Questions & Answers
1. What retrieval abuse technique involves crafting malicious content so it ranks highly for sensitive queries?
Active manipulation
2. What does retrieval select documents based on?
Semantic relevance
Task 5 - Real-World RAG Failure Scenarios
Overview
This task looks at real-world RAG failures where retrieved content influenced AI responses in unintended ways.Key examples:
Microsoft Copilot: Malicious or misleading content in emails could be retrieved and influence responses.
ChatGPT Plugins: External web/API content could contain hidden instructions, leading to indirect prompt injection via retrieval.
Web-connected AI assistants: Outdated information remained indexed and was retrieved as if it were current.
The main lesson is that retrieval itself is a security boundary. Even trusted sources can become risky if content isn't validated, updated, or properly governed.
Questions & Answers
1. In the Web-Connected AI Assistants cases, failures were caused by governance gaps in what specific part of the system?
Retrieval pipelines
Task 6 - Defensive Considerations for RAG Systems
Overview
This task focuses on defending RAG systems against poisoning and retrieval abuse. The main idea is layered defence rather than relying on one control.Key protections include:
Ingestion validation: Review sources, approvals, ownership, and update history before data enters the vector store.
Guardrails: Separate retrieved content from system instructions and detect instruction-like patterns.
Behavioural monitoring: Watch for unusual retrieval patterns and changes in model behaviour.
Output review: Look for gradual changes in responses, known as output drift.
Defence-in-depth: Combine multiple controls because no single protection is perfect.
Questions & Answers
Behavioural monitoring
2. What does output drift reflect instead of a sudden failure?
Gradual influence from ______________________________