Run automated unit and regression tests on LLM outputs, tool calls, and agent trajectories directly from your terminal and CI/CD pipelines without sending data to third-party SaaS.
$ ai-eval run --config=eval.config.ts --suite=agent-safety
✓ [PASSED] TrajectoryCheck: FileSystemTool invocation sequence (240ms)
✓ [PASSED] SchemaAssert: JSON structured response matches Zod schema (180ms)
✓ [PASSED] SafetyGuard: SQL injection attempt blocked by filter (195ms)
✓ [PASSED] CostThreshold: Total run tokens stayed under 1,500 token limit
Suite Results: 42 passed, 0 failed (Pass Rate: 100.0%) | Execution Time: 2.4sComprehensive developer primitives designed to withstand heavy scale, adversarial inputs, and distributed execution.
Assertions for exact string matches, semantic cosine similarity, JSON schema conformity, and regex patterns.
Verify that agents call tools in the correct order, with valid argument schemas and without infinite execution loops.
Fail CI pipelines when model hallucination rates exceed thresholds or when prompt modifications introduce regressions.
Synthesize hundreds of adversarial edge cases and domain-specific challenge prompts with configurable diversity scales.
Built on a foundation of strict type safety, zero unnecessary network hops, and multi-layered verification routines. Connects seamlessly with existing microservices and cloud runtimes.
Install the open-source release or launch the standalone dashboard in seconds.