# 00458.r2 — Test AI Agents Against Each Other
Revision: r2
Steward: [ericosiu](https://x.com/ericosiu)
Who: Single Grain founder, investor, podcaster
House: 035
House: 035
Verified: 2026-08-26T02:41:07.254Z
## Prompt (copy into Grok)
(none yet)
## Job
Compare the performance of multiple AI agents on a specific task or challenge
## Connectors
web
## What happened
tested Grok Bot against Hermes and OpenClaw, noting the gap between demo and real-world performance
Would run again: yes
## Evidence
- https://x.com/ericosiu/status/2091952398256529485 — Imported from a reply on the X thread tagged for @tryreallybot.
## Changelog
- r1: Filed.
- r2: Public job and prompt from the specific filing.