TryHackMe Walkthrough: Securing AI Systems

Task 1 - Introduction

This task introduces TryAssist, an AI-powered code review assistant connected to internal documentation, repositories, APIs, and CI/CD.

The main lesson is that adding AI significantly expands an organisation's attack surface. Security teams need to consider not only the LLM, but also the tools, data, permissions, logging, and integrations surrounding it.

Key Takeaway

AI security starts with understanding the entire architecture, not just the model.

Task 2 - Anatomy of an AI System

This task explores the components of a production AI system, including the API Gateway, orchestration layer, prompt construction, LLM, tool layer, vector store, and logging.

It also introduces trust boundaries, showing how data moves between users, the AI model, external data sources, and tools.

Key Takeaway

Every connection between the LLM and another system can introduce a new attack surface.

Questions & Answers

Q1. What layer combines the system prompt, user input, and retrieved context?

Answer: Prompt Construction

Q2. What boundary does LLM output cross when it triggers a database query?

Answer: LLM-to-tools


Task 3 - The AI Attack Surface

This task introduces the main frameworks used to classify and manage AI security risks:

  • OWASP LLM Top 10

  • MITRE ATLAS

  • NIST AI RMF

It focuses on risks such as prompt injection, sensitive information disclosure, excessive agency, system prompt leakage, improper output handling, and unbounded consumption.

Key Takeaway

AI security benefits from combining application-security frameworks with AI-specific threat models.

Questions & Answers

Q1. Which OWASP category covers LLM output being used to execute SQL injection against a backend database?

Answer: LLM05

Q2. What is the MITRE knowledge base designed for adversary tactics and techniques against AI/ML systems?

Answer: ATLAS


Task 4 - System-Level Threats

This task focuses on five important AI security risks:

  • Improper Output Handling

  • Excessive Agency

  • System Prompt Leakage

  • Unbounded Consumption

  • Sensitive Information Disclosure

The key security principles are least privilege, output validation, rate limiting, data protection, and avoiding secrets in system prompts.

Key Takeaway

AI output should never automatically be treated as trusted, especially when it can trigger actions in other systems.

Questions & Answers

Q1. The Air Canada chatbot incident is classified under which category?

Answer: LLM09

Q2. What are the three dimensions of excessive agency?

Answer: Excessive Functionality, Excessive Permissions, Excessive Autonomy

Q3. Extracting internal API endpoints from a system prompt falls under which category?

Answer: LLM07

Q4. Thousands of maximum-length requests generating a large bill fall under which category?

Answer: LLM10


Task 5 - Secure Design Patterns

This task covers practical ways to secure AI architectures.

Important controls include:

  • Defence in depth

  • Least privilege

  • Human approval for high-risk actions

  • Input and output validation

  • Monitoring and observability

  • MLSecOps

Key Takeaway

AI security should be built into the system from the design stage rather than added after deployment.

Questions & Answers

Q1. What principle states that every AI component should have the minimum permissions required?

Answer: Least Privilege

Q2. What practice integrates security into the machine-learning lifecycle?

Answer: MLSecOps


Task 6 - Auditing TryAssist

This task puts the concepts into practice by auditing the TryAssist AI assistant.

The audit looks at its:

  • Capabilities

  • Permissions

  • Autonomy

  • System instructions

  • Data retention

The assessment identifies excessive permissions and autonomous capabilities that could create significant security risks.

Key Takeaway

An AI assistant can become a high-risk system when it has excessive permissions or the ability to perform actions without human approval.

Questions & Answers

Q1. What action does TryAssist take automatically without human approval?

Answer: Merge Pull Requests

Q2. What database role does TryAssist report operate under?

Answer: db_admin

Q3. What security control is missing from conversation logging?

Answer: PII Filtering


Task 7 - Conclusion

The final task brings the concepts together.

The main lesson is that securing AI requires looking at the whole system:

User > Application > LLM > Tools > Data > External Systems

Security teams need to understand what the AI can access, what actions it can perform, and what could happen if the model is manipulated.

Key Takeaway

Don't just ask whether the AI model is secure. Ask what the AI can access, what it can do, and what happens if it is compromised.

Popular posts from this blog

TryHackMe Walkthrough: AI Security Threats