TryHackMe Walkthrough: Agent Evaluation

Task 1 - Introduction

This section introduces the purpose of evaluating an AI security agent and explains why a correct final answer does not necessarily mean the investigation was performed correctly. The agent needs to be evaluated based on the evidence it retrieves, the decisions it makes, and whether changes improve its overall behaviour.

Task 2 - Establish the Baseline

Before improving the agent, you first need to understand how it behaves without any changes. This starting point is called a baseline, and it provides a reference for measuring whether later improvements actually make the agent more reliable.

In this task, you will use the NorthStar Agent Evaluation Workspace to run the Security Investigation Agent against five reviewed security cases and record its initial performance.

Questions & Answers

Q1. How many reviewed cases passed the baseline evaluation?

Answer: 3

Q2. Which search-related measurement failed for the New Device investigation?

Answer: Query Target

Task 3 - Diagnose the Failure

The baseline showed that the New Device investigation failed two measurements:

Query Target Fail
Verdict Fail

A failed score tells us what went wrong, but not necessarily why. In this task, you will inspect the investigation evidence to identify the cause of the incorrect search target. Return to the NorthStar Agent Evaluation Workspace and click Investigate Failure.

Questions & Answers

Q1. Which SIEM field should be used for a platform device identifier?

Answer: device.device_id

Q2. Does one zero-result search prove that related activity is absent? (Yea / Nay)

Answer: Nay

Task 4 - Improve the Candidate

In the previous task, you diagnosed the New Device failure.

The agent used valid SIEM syntax, but searched the platform identifier using the wrong field:

The search returned no results, and the investigation stopped too early. Now that the cause is clear, you can make a focused improvement that addresses the specific failure rather than changing the agent blindly. Return to the NorthStar Agent Evaluation Workspace and continue to the improvement step.

Questions & Answers

Q1. Which verdict applies when reliable evidence contradicts a condition asserted by the alert?

Answer: FalsePositive

Q2. Did fixing the evidence search automatically fix the verdict? (Yea / Nay)

Answer: Nay

Task 5 - Test for Regressions

The targeted evaluation now passes:

✓ Credential attack
✓ Unusual location
✓ New device

3 / 3 passed

This confirms that the changes improved the investigations they were intended to fix, but the candidate is not ready yet. An improvement can correct one behaviour while unintentionally breaking another that previously worked; this is known as a regression.

For example, stronger FalsePositive guidance might fix the New Device case while incorrectly changing a routine office sign-in that should remain BenignPositive. Once the full regression reaches 5 / 5, the workspace will reveal the room flag.

Questions & Answers

Q1. What type of testing checks whether a change broke behaviour that previously worked?

Answer: Regression

Q2. How many reviewed cases passed the final regression?

Answer: 5 

Q3. What is the flag revealed after completing the full regression?

Answer: THM{TRACE_TEST_REDACTED}

Popular posts from this blog

TryHackMe Walkthrough: AI Security Threats