Guidance and xUnit samples for evaluating .NET AI agents with AgentEval — stochastic tests, red-team scanning, trace replay, memory benchmarks, and CI templates
-
Updated
May 20, 2026
Guidance and xUnit samples for evaluating .NET AI agents with AgentEval — stochastic tests, red-team scanning, trace replay, memory benchmarks, and CI templates
AgentEval (AutoGen 0.4) Sample Implementation
FastMCP server integrating a stateful app into RL sandbox environments — populate/snapshot/restore hooks with content-addressed state digests, Docker packaging, and a 51-test suite spanning the app, state, and protocol layers.
OpenEnv benchmark where AI agents detect hallucinations from stale documents, identify misleading sources, repair the knowledge base, and verify corrected answers through a scored multi-step RAG environment.
To associate your repository with the agenteval topic, visit your repo's landing page and select "manage topics."