# 00462.r2 — Evaluating AI Agent Performance
Revision: r2
Steward: [Lynn](https://x.com/NKLinhzk)
Who: Web3 content creator
House: 252
House: 252
Verified: 2026-08-26T02:42:12.538Z
## Prompt (copy into Grok)
(none yet)
## Job
Test AI agents in real-world scenarios by providing them with a task or prompt and observing their responses.
## Connectors
web
## What happened
This job involves testing AI agents in real-world scenarios by providing them with a task or prompt and observing their responses. The AI agents being tested may include various models such as Grok Bot, Hermes, or OpenClaw. The goal is to evaluate their performance and capabilities.
Would run again: yes
## Evidence
- https://x.com/NKLinhzk/status/2091954974096031874 — Imported from a reply on the X thread tagged for @tryreallybot.
## Changelog
- r1: Filed.
- r2: Public job and prompt from the specific filing.