TryHackMe Walkthrough: AI System Reconnaissance

TryHackMe Walkthrough: AI Reconnaissance

AI systems are becoming part of modern infrastructure, but finding them is not always as straightforward as scanning for traditional services. AI platforms introduce new ports, APIs, protocols, model-serving endpoints, experiment trackers, vector databases, and supporting services that may not be correctly identified by standard security tools.

This TryHackMe room focuses on AI reconnaissance, finding, identifying, and enumerating AI infrastructure exposed within a network.

Instead of starting with exploitation, the focus is on understanding what is actually deployed.


Task 1 - Introduction

This task introduces the concept of AI reconnaissance and explains how AI infrastructure differs from a traditional network.

What is AI Reconnaissance?

AI reconnaissance is the process of discovering AI and ML components in an environment, identifying what technologies they use, and determining what information or functionality they expose.

The important point is that the goal is not immediately to exploit the service. The first step is understanding the AI infrastructure that exists.

One important lesson from this task is that AI infrastructure can significantly expand the network attack surface. An organisation may have many additional services that would not normally appear during a traditional web application assessment.

Why AI Reconnaissance Matters

Exposed AI services can reveal useful information about an organisation's environment.

🔑 Key Takeaway

AI reconnaissance starts with knowing what to look for. Ports, protocols, API paths, and service responses can reveal AI infrastructure that a normal network scan may not immediately identify.


Task 2 - Discovering AI Infrastructure

The goal of this task is to identify AI services running inside the Cyphira internal network.

Step 1 - Scan for AI Services

Step 2 - Scan Traditional Services

Step 3 - Map Ports to AI Components

Using the reference table from the room, we can associate ports with likely technologies.

This is why having an AI-specific port reference is useful. A normal scan might simply tell you that port 8000 is open. AI reconnaissance asks the next question: what AI service is actually running there?

Questions & Answers

Q1. What is the IP address of the host running an HTTP service on port 8888 in your scan results?

Answer: 10.10.45.20

Port 8888 is commonly associated with Jupyter Notebook. The room later confirms this host by probing:

curl http://10.10.45.20:8888/api/kernels

Q2. Which port does MLflow Tracking Server run on by default?

Answer: 5000


Task 3 - Fingerprinting AI Services

Standard Nmap service detection does not always correctly identify AI services. For example, a model server running on port 8000 may simply appear as an HTTP service.

🔑 Key Takeaways

This room demonstrates that AI reconnaissance is different from traditional network reconnaissance.

  • AI infrastructure introduces its own set of ports and services.
  • Knowing the common AI ports makes discovery much easier.
  • Nmap alone may not correctly identify AI frameworks.
  • HTTP headers can provide strong framework fingerprints.
  • JSON response structures can reveal model-serving technologies.
  • Error messages can expose framework-specific information.
  • gRPC reflection can reveal an AI service's API structure.
  • Vector databases and notebooks can expose additional information about an AI environment.

The main lesson is simple: finding an AI service is only the first step. Fingerprinting tells you what you actually found.

Questions & Answers

Q1. Which unique HTTP response header does the service on 10.10.45.15:8000 return to identify as an NVIDIA product?

Answer: NV-Status

This is an indicator of NVIDIA Triton Inference Server.

Q2. When you run grpcurl against 10.10.45.15:8001, what is the name of the inference service listed in the reflection output?

Answer: inference.GRPCInferenceService


Task 4 - Enumerating AI Systems

Fingerprinting tells us what a service is. Enumeration tells us what information the service exposes.

This is where the reconnaissance starts becoming much more useful.

MLflow Enumeration

MLflow is one of the most valuable services to enumerate because it can contain information about experiments, models, training runs, and artifact storage.

Step 1 - List Experiments

Step 2 - List Registered Models

Step 3 - Get Model Version Details

Step 4 - Search Training Runs

Step 5 - List Artifacts

🔑 Key Takeaway

Fingerprinting tells you what AI service is running, but enumeration tells you what is inside that service.

With MLflow, Triton, vector databases and Jupyter, exposed APIs can reveal models, versions, artifacts, configurations, metadata and even credentials. This makes unauthenticated AI management interfaces a valuable source of information during reconnaissance.

The main lesson from this task is simple: once an AI service is identified, always check what its APIs expose. Metadata that looks harmless individually can reveal enough information to map the underlying AI environment.

Questions & Answers

Q1. What MLflow REST API endpoint would you use to retrieve the artifact storage location for a specific model version?

Answer: /api/2.0/mlflow/model-versions/search

The model version response contains the source field, which can reveal the artifact URI.

Q2. What is the cleartext password for the MLflow service account stored in the Jupyter notebook on 10.10.45.20?

Answer: Cyphira-MLfl0w-2024!.


Task 5 - Mapping the AI Attack Surface

Finding an individual exposed service is useful, but an attack surface map shows how those services interact.

How AI Expands the Attack Surface

An AI environment can contain many interconnected services:

  • Inference servers communicate with vector databases.
  • Orchestration platforms manage model deployments.
  • Jupyter notebooks connect to MLflow and cloud services.
  • Prometheus collects metrics from model servers.
  • Model registries point to cloud storage containing model artifacts.

This means the security boundary does not necessarily stop at the external firewall.

🔑 Key Takeaway

  • AI infrastructure often has multiple interconnected components, so a weakness in one service can expose other parts of the environment.
  • MLflow, Kubeflow and TorchServe can introduce significant exposure when authentication or management interfaces are misconfigured.
  • Model registries contain valuable information such as model names, versions, artifact locations, run IDs and user IDs.
  • Supply-chain reconnaissance can reveal Hugging Face tokens, internal packages and model download sources.
  • The findings can be mapped to MITRE ATLAS techniques such as Active Scanning, Discover ML Artifacts and ML Supply Chain Compromise.
  • The ShadowRay case study demonstrates how reconnaissance can progress from discovery → fingerprinting → enumeration → compromise.

In short: AI reconnaissance is not just about finding open ports. The goal is to understand how the AI components connect, what information they expose, and where those connections could create opportunities for further attack.

Questions & Answers

Q1. The Cyphira Jupyter notebook contains a Hugging Face token, and the internal-kb-embedder model references sentence-transformers/all-MiniLM-L6-v2 as its base model. What ATLAS technique ID covers the risk of these exposed supply chain dependencies?

Answer: AML.T0010 - ML Supply Chain Compromise

Q2. You scanned the Cyphira subnet with Nmap, probed endpoints with curl, and extracted metadata from MLflow APIs. All of these activities fall under one overarching ATLAS tactic. What is its ID?

Answer: AML.TA0002 - Reconnaissance


Task 6 – Structured Reconnaissance Methodology and Detection

This Task brings everything together into a repeatable 5-phase AI reconnaissance methodology. It also changes perspective and looks at what all of this reconnaissance activity looks like from the defender's side in SIEM logs.

🔑 KeyTakeaway

  • AI reconnaissance can be performed as a repeatable five-phase process.
  • Passive reconnaissance can reveal AI infrastructure before touching the target.
  • AI-specific ports make active scanning more effective.
  • API fingerprinting helps identify the framework behind exposed services.
  • Metadata extraction reveals models, artifacts, users and deployment information.
  • Supply-chain review can expose tokens, model sources and dependency weaknesses.
  • The same reconnaissance activity creates detectable SIEM patterns.
  • Authentication, network restrictions, scoped tokens and restricted metrics can significantly reduce the exposed reconnaissance surface.

Questions & Answers

Q1. A SIEM log shows requests to /api/2.0/mlflow/registered-models/list from an IP with no corresponding MLflow UI session. What tool's access pattern does this match?

Answer: MLOKit

The room identifies raw MLflow API requests without the normal UI session pattern as a characteristic of scripted enumeration, specifically matching the access pattern produced by MLOKit.

Q2. What is the single most effective quick win for preventing unauthenticated access to the MLflow tracking server?

Answer: Enable MLflow authentication

The task recommends enabling authentication with MLFLOW_TRACKING_USERNAME and MLFLOW_TRACKING_PASSWORD, or deploying MLflow behind an authenticating reverse proxy.

Popular posts from this blog

TryHackMe Walkthrough: AI Security Threats