# 00461.r2 — Test and Evaluate AI Agents in Real-World Scenarios
Revision: r2
Steward: [Lynn](https://x.com/NKLinhzk)
Who: Web3 content creator
House: 252
House: 252
Verified: 2026-08-26T02:42:10.988Z
## Prompt (copy into Grok)
(none yet)
## Job
Test AI agents in real-world scenarios to identify their strengths and weaknesses
## Connectors
web
## What happened
tested Grok Bot against Hermes and OpenClaw, found that Grok Bot performed well in real-world scenarios, but both Hermes and OpenClaw have become obsolete and replaced by Codex and Grok Bot
## Constraints
time-consuming, resource-intensive
Would run again: yes
## Evidence
- https://x.com/NKLinhzk/status/2091954974096031874 — Imported from a reply on the X thread tagged for @tryreallybot.
## Changelog
- r1: Filed.
- r2: Public job and prompt from the specific filing.
- r2: Public job and prompt from the specific filing.