
About
Most agent evaluations assess only the final response. Redline evaluates the agent's actual behavior by examining the server's internal record of its actions, not the agent's own summary. A single task is executed against Claude Code, Codex, and the agent from your own repository, each within an isolated, identical container. This is supplemented by 16 security packs containing 11,204 adversarial test cases, which are delivered as real attacks would be, concealed within elements like a support ticket, a document footer, or a tool output. Your agent connects to the Redline development environment and its operations remain entirely on your local machine.
Launched
August 25, 2026Week 25
Builder
BU
BuilderComments
Sign in to leave a comment
Sign In