TryHackMe Walkthrough: Agent Evaluation
Task 1 - Introduction
This section introduces the purpose of evaluating an AI security agent and explains why a correct final answer does not necessarily mean the investigation was performed correctly. The agent needs to be evaluated based on the evidence it retrieves, the decisions it makes, and whether changes improve its overall behaviour.
Task 2 - Establish the Baseline
Before improving the agent, you first need to understand how it behaves without any changes. This starting point is called a baseline, and it provides a reference for measuring whether later improvements actually make the agent more reliable.
In this task, you will use the NorthStar Agent Evaluation Workspace to run the Security Investigation Agent against five reviewed security cases and record its initial performance.
Questions & Answers
Q1. How many reviewed cases passed the baseline evaluation?
Answer: 3
Q2. Which search-related measurement failed for the New Device investigation?
Answer: Query Target
Task 3 - Diagnose the Failure
The baseline showed that the New Device investigation failed two measurements:
Query Target Fail
Verdict Fail
A failed score tells us what went wrong, but not necessarily why. In this task, you will inspect the investigation evidence to identify the cause of the incorrect search target. Return to the NorthStar Agent Evaluation Workspace and click Investigate Failure.
Questions & Answers
Q1. Which SIEM field should be used for a platform device identifier?
Answer: device.device_id
Q2. Does one zero-result search prove that related activity is absent? (Yea / Nay)
Answer: Nay
Task 4 - Improve the Candidate
In the previous task, you diagnosed the New Device failure.
The agent used valid SIEM syntax, but searched the platform identifier using the wrong field:
The search returned no results, and the investigation stopped too early. Now that the cause is clear, you can make a focused improvement that addresses the specific failure rather than changing the agent blindly. Return to the NorthStar Agent Evaluation Workspace and continue to the improvement step.
Questions & Answers
Q1. Which verdict applies when reliable evidence contradicts a condition asserted by the alert?
Answer: FalsePositive
Q2. Did fixing the evidence search automatically fix the verdict? (Yea / Nay)
Answer: Nay
Task 5 - Test for Regressions
The targeted evaluation now passes:
✓ Credential attack
✓ Unusual location
✓ New device
3 / 3 passed
This confirms that the changes improved the investigations they were intended to fix, but the candidate is not ready yet. An improvement can correct one behaviour while unintentionally breaking another that previously worked; this is known as a regression.
For example, stronger FalsePositive guidance might fix the
New Device case while incorrectly changing a routine office
sign-in that should remain BenignPositive. Once the full regression reaches 5 / 5, the workspace will reveal
the room flag.
Questions & Answers
Q1. What type of testing checks whether a change broke behaviour that previously worked?
Answer: Regression
Q2. How many reviewed cases passed the final regression?
Answer: 5
Q3. What is the flag revealed after completing the full regression?
Answer: THM{TRACE_TEST_REDACTED}