TryHackMe Walkthrough: AI System Reconnaissance
TryHackMe Walkthrough: AI Reconnaissance
AI systems are becoming part of modern infrastructure, but finding them is not always as straightforward as scanning for traditional services. AI platforms introduce new ports, APIs, protocols, model-serving endpoints, experiment trackers, vector databases, and supporting services that may not be correctly identified by standard security tools.
This TryHackMe room focuses on AI reconnaissance, finding, identifying, and enumerating AI infrastructure exposed within a network.
Instead of starting with exploitation, the focus is on understanding what is actually deployed.
Task 1 - Introduction
This task introduces the concept of AI reconnaissance and explains how AI infrastructure differs from a traditional network.
What is AI Reconnaissance?
AI reconnaissance is the process of discovering AI and ML components in an environment, identifying what technologies they use, and determining what information or functionality they expose.
The important point is that the goal is not immediately to exploit the service. The first step is understanding the AI infrastructure that exists.
One important lesson from this task is that AI infrastructure can significantly expand the network attack surface. An organisation may have many additional services that would not normally appear during a traditional web application assessment.
Why AI Reconnaissance Matters
Exposed AI services can reveal useful information about an organisation's environment.
🔑 Key Takeaway
AI reconnaissance starts with knowing what to look for. Ports, protocols, API paths, and service responses can reveal AI infrastructure that a normal network scan may not immediately identify.
Task 2 - Discovering AI Infrastructure
The goal of this task is to identify AI services running inside the Cyphira internal network.
Step 1 - Scan for AI Services
Step 2 - Scan Traditional Services
Step 3 - Map Ports to AI Components
Using the reference table from the room, we can associate ports with likely technologies.
This is why having an AI-specific port reference is useful. A normal scan might simply tell you that port 8000 is open. AI reconnaissance asks the next question: what AI service is actually running there?
Questions & Answers
Q1. What is the IP address of the host running an HTTP service on port 8888 in your scan results?
Answer: 10.10.45.20
Port 8888 is commonly associated with Jupyter Notebook. The room later confirms this host by probing:
curl http://10.10.45.20:8888/api/kernels
Q2. Which port does MLflow Tracking Server run on by default?
Answer: 5000
Task 3 - Fingerprinting AI Services
Standard Nmap service detection does not always correctly identify AI services. For example, a model server running on port 8000 may simply appear as an HTTP service.
🔑 Key Takeaways
This room demonstrates that AI reconnaissance is different from traditional network reconnaissance.
- AI infrastructure introduces its own set of ports and services.
- Knowing the common AI ports makes discovery much easier.
- Nmap alone may not correctly identify AI frameworks.
- HTTP headers can provide strong framework fingerprints.
- JSON response structures can reveal model-serving technologies.
- Error messages can expose framework-specific information.
- gRPC reflection can reveal an AI service's API structure.
- Vector databases and notebooks can expose additional information about an AI environment.
The main lesson is simple: finding an AI service is only the first step. Fingerprinting tells you what you actually found.
Questions & Answers
Q1. Which unique HTTP response header does the service on 10.10.45.15:8000 return to identify as an NVIDIA product?
Answer: NV-Status
This is an indicator of NVIDIA Triton Inference Server.
Q2. When you run grpcurl against 10.10.45.15:8001, what is the name of the inference service listed in the reflection output?
Answer: inference.GRPCInferenceService
Task 4 - Enumerating AI Systems
Fingerprinting tells us what a service is. Enumeration tells us what information the service exposes.
This is where the reconnaissance starts becoming much more useful.
MLflow Enumeration
MLflow is one of the most valuable services to enumerate because it can contain information about experiments, models, training runs, and artifact storage.
Step 1 - List Experiments
Step 2 - List Registered Models
Step 3 - Get Model Version Details
Step 4 - Search Training Runs
Step 5 - List Artifacts
🔑 Key Takeaway
Fingerprinting tells you what AI service is running, but enumeration tells you what is inside that service.
With MLflow, Triton, vector databases and Jupyter, exposed APIs can reveal models, versions, artifacts, configurations, metadata and even credentials. This makes unauthenticated AI management interfaces a valuable source of information during reconnaissance.
The main lesson from this task is simple: once an AI service is identified, always check what its APIs expose. Metadata that looks harmless individually can reveal enough information to map the underlying AI environment.
Questions & Answers
Q1. What MLflow REST API endpoint would you use to retrieve the artifact storage location for a specific model version?
Answer: /api/2.0/mlflow/model-versions/search
The model version response contains the source field, which can reveal the artifact URI.
Q2. What is the cleartext password for the MLflow service account stored in the Jupyter notebook on 10.10.45.20?
Answer: Cyphira-MLfl0w-2024!.
Task 5 - Mapping the AI Attack Surface
Finding an individual exposed service is useful, but an attack surface map shows how those services interact.
How AI Expands the Attack Surface
An AI environment can contain many interconnected services:
- Inference servers communicate with vector databases.
- Orchestration platforms manage model deployments.
- Jupyter notebooks connect to MLflow and cloud services.
- Prometheus collects metrics from model servers.
- Model registries point to cloud storage containing model artifacts.
This means the security boundary does not necessarily stop at the external firewall.
🔑 Key Takeaway
- AI infrastructure often has multiple interconnected components, so a weakness in one service can expose other parts of the environment.
- MLflow, Kubeflow and TorchServe can introduce significant exposure when authentication or management interfaces are misconfigured.
- Model registries contain valuable information such as model names, versions, artifact locations, run IDs and user IDs.
- Supply-chain reconnaissance can reveal Hugging Face tokens, internal packages and model download sources.
- The findings can be mapped to MITRE ATLAS techniques such as Active Scanning, Discover ML Artifacts and ML Supply Chain Compromise.
- The ShadowRay case study demonstrates how reconnaissance can progress from discovery → fingerprinting → enumeration → compromise.
In short: AI reconnaissance is not just about finding open ports. The goal is to understand how the AI components connect, what information they expose, and where those connections could create opportunities for further attack.
Questions & Answers
Q1. The Cyphira Jupyter notebook contains a Hugging Face token, and the internal-kb-embedder model references sentence-transformers/all-MiniLM-L6-v2 as its base model. What ATLAS technique ID covers the risk of these exposed supply chain dependencies?
Answer: AML.T0010 - ML Supply Chain Compromise
Q2. You scanned the Cyphira subnet with Nmap, probed endpoints with curl, and extracted metadata from MLflow APIs. All of these activities fall under one overarching ATLAS tactic. What is its ID?
Answer: AML.TA0002 - Reconnaissance
Task 6 – Structured Reconnaissance Methodology and Detection
This Task brings everything together into a repeatable 5-phase AI reconnaissance methodology. It also changes perspective and looks at what all of this reconnaissance activity looks like from the defender's side in SIEM logs.
🔑 KeyTakeaway
- AI reconnaissance can be performed as a repeatable five-phase process.
- Passive reconnaissance can reveal AI infrastructure before touching the target.
- AI-specific ports make active scanning more effective.
- API fingerprinting helps identify the framework behind exposed services.
- Metadata extraction reveals models, artifacts, users and deployment information.
- Supply-chain review can expose tokens, model sources and dependency weaknesses.
- The same reconnaissance activity creates detectable SIEM patterns.
- Authentication, network restrictions, scoped tokens and restricted metrics can significantly reduce the exposed reconnaissance surface.
Questions & Answers
Q1. A SIEM log shows requests to /api/2.0/mlflow/registered-models/list from an IP with no corresponding MLflow UI session. What tool's access pattern does this match?
Answer: MLOKit
The room identifies raw MLflow API requests without the normal UI session pattern as a characteristic of scripted enumeration, specifically matching the access pattern produced by MLOKit.
Q2. What is the single most effective quick win for preventing unauthenticated access to the MLflow tracking server?
Answer: Enable MLflow authentication
The task recommends enabling authentication with MLFLOW_TRACKING_USERNAME and MLFLOW_TRACKING_PASSWORD, or deploying MLflow behind an authenticating reverse proxy.