Decision orchestration and reconciliation for AI changes.
-
Updated
Mar 30, 2026 - Rust
Decision orchestration and reconciliation for AI changes.
An graph-eval framework for LLM's
In this we evaluate the LLM responses and find accuracy
Semantic testing for Microsoft Copilot Studio agents using Pytest and DeepEval G-Eval metrics (LLM-as-a-Judge). Generates interactive HTML reports for agent response quality.
Eval-driven model router: pick the best model per task, score with multi-metric eval, track win-rates on a Postgres scoreboard.
To associate your repository with the geval topic, visit your repo's landing page and select "manage topics."