TryHackMe Walkthrough: AI Security Threats

 Task 1 - Introduction

This section introduces the focus of the room and explains why AI security is becoming increasingly important. Rather than covering how AI works, it shifts the discussion toward the security risks that AI systems introduce, how attackers are using AI to strengthen their attacks, and how defenders can leverage the same technology to improve detection and response. It also highlights the importance of securely adopting AI using established security practices and frameworks.


Questions & Answers

Q1. I'm ready to learn about AI/ML security threats!

Answer:
No answer required (just complete the task).

Task 2 - Vulnerabilities in AI Models

This section introduces the security risks that are unique to AI systems. Unlike traditional software vulnerabilities, these weaknesses arise from how AI models are trained, deployed, and interact with users. It also introduces the MITRE ATLAS framework, which categorizes adversarial tactics and techniques targeting AI systems. Understanding these vulnerabilities helps security professionals identify potential attack vectors and implement appropriate safeguards when deploying AI applications.

Key AI Vulnerabilities

  • Prompt Injection – Malicious prompts manipulate or override a model's original instructions, causing unintended behaviour or disclosure of sensitive information.

  • Data Poisoning - Attackers tamper with training data to influence a model's behaviour, resulting in inaccurate or biased outputs.

  • Model Theft - An attacker replicates or steals a model, often by repeatedly querying its API to build a functional clone.

  • Privacy Leakage - A model unintentionally exposes sensitive information that was present in its training data.

  • Model Drift - A model's accuracy decreases over time as real-world data and conditions change, requiring continuous monitoring and retraining.

The practical exercise demonstrates prompt injection, where the objective is to manipulate an AI assistant into revealing its hidden system prompt, highlighting how improperly secured LLMs can expose confidential information.


Questions & Answers

Q1. What MITRE framework was developed specifically to map tactics and techniques used against AI systems?

Answer: ATLAS

Q2. What AI vulnerability occurs when user input overrides the original instructions provided to a model?

Answer: Prompt injection

Q3. What attack involves manipulating training data to cause a model to produce incorrect or biased outputs?

Answer: Data poisoning

Q4. What attack involves repeatedly querying a model's API to train a clone that replicates its behaviour?

Answer: Model theft

Q5. What term describes the gradual degradation of a model's performance as the environment it was trained on changes over time?

Answer: Model drift

Q6. What's the flag?

Answer: T

Task 3 - AI-Enhanced Attacks

This section focuses on how attackers are leveraging AI to improve traditional cyberattacks. Rather than introducing entirely new attack methods, AI enhances existing techniques by making them more convincing, scalable, and accessible. As a result, attacks that once required significant technical expertise can now be generated quickly using AI-powered tools.

AI-Enhanced Attack Techniques

  • AI-Generated Malware - Attackers use generative AI to create or modify malicious code rapidly, lowering the barrier to malware development.

  • Deepfakes - AI generates realistic voices, images, or videos that impersonate trusted individuals, making fraud and identity-based attacks more convincing.

  • AI-Enhanced Phishing - Large language models can produce highly personalized and grammatically correct phishing emails, making them much harder to distinguish from legitimate communications.

In the practical exercise, the objective is to review messages in a simulated inbox, identify the AI-powered attack technique being used, and explain how AI contributed to each attack.


Questions & Answers

Q1. What AI technique is used to generate convincing replicas of a person's voice or appearance?

Answer: Deepfakes

Q2. What common initial access method has become significantly harder to detect due to AI's ability to generate fluent, targeted content at scale?

Answer: Phishing

Q3. What is the flag?

Answer: 

Task 4 - Defensive AI

This section explores how AI can strengthen cybersecurity by improving the speed and efficiency of security operations. Instead of viewing AI solely as a threat, it highlights how defenders use AI to detect attacks earlier, automate repetitive tasks, and support analysts during investigations. AI enables security teams to process vast amounts of data in seconds, helping organizations reduce response times and minimize the impact of security incidents.

Key Defensive AI Capabilities

  • Analysis - AI identifies anomalies and suspicious activity across logs, network traffic, and endpoint telemetry much faster than manual analysis.

  • Prediction - Machine learning models detect potential threats, such as phishing emails or malicious behaviour, before they cause damage.

  • Summarisation - Large Language Models (LLMs) condense lengthy incident reports, alerts, and threat intelligence into concise summaries for faster decision-making.

  • Investigation - AI assists analysts by interpreting logs, suggesting investigative steps, supporting threat hunting, and helping determine the root cause of incidents.

The practical exercise demonstrates these capabilities by using an AI assistant to analyse a firewall log, triage a phishing email, summarise the incident for leadership, and suggest additional threats that may still exist in the environment.


Questions & Answers

Q1. According to IBM, how many days faster does AI help identify and contain breaches?

Answer: 108

Q2. What Microsoft product is mentioned as an example of a security tool leveraging AI for analysis?

Answer: Microsoft Defender for Endpoint

Q3. What defensive AI capability involves feeding an LLM raw logs to help identify what happened during a security incident?

Answer: Investigation

Q4. What's the flag?

Answer: