Eval framework. Define correct, test against it, get results.
-
Updated
Feb 17, 2026 - Go
Eval framework. Define correct, test against it, get results.
Find what your AI agent gets wrong — before you have a rubric. Qualitative eval for PMs.
RAG context precision and context recall evaluator measuring retrieved chunk relevance and rank quality
RAG context precision and context recall evaluator measuring retrieved chunk relevance and rank quality
A web-based interactive demo for the GuessArena evaluation framework
Curated AI agent evaluation skills from Microsoft's Eval Guide — plan, generate, run, and interpret eval suites for Copilot Studio agents
One-stop CLI for running structured skill evals across all agent harnesses
4-model parallel planning workflow with eval framework — Claude, Gemini, Codex, GLM-5 · OpenClaw ecosystem
Binary safety verdicts (SAFE/HELD/LEAK/MISS/BROKE) + persona fan-out for LLM pipeline evals
Open-source evaluation framework for MCP servers powered by Claude. Auto-discovers tools, generates test scenarios, runs LLM-as-judge scoring to assess correctness and safety, audits for security issues, and visualizes traces in a web dashboard.
Open-source evaluation framework for LLM agents. Run head-to-head A/B tests, score with LLM-as-judge rubrics, and visualize results in a Streamlit dashboard. Model-agnostic, self-hosted, zero external infrastructure.
Observability layer for multi-step AI pipelines — traces execution, auto-diagnoses root causes, and builds a growing eval dataset from human feedback
Evaluation framework for testing LLM outputs locally. Define prompt templates and custom scorers, run evals against OpenAI or Anthropic models, store results in SQLite, and browse via Rich CLI dashboard or FastAPI UI. Lightweight, self-contained, extensible Python library.
Python tool for extracting validated, schema-defined JSON from documents (PDF, HTML, Markdown, plaintext) using LLMs, with source grounding and a built-in eval harness. Works with Anthropic Claude or OpenAI.
Self-hosted evaluation framework for tool-calling LLM agents. Define tasks in YAML with expected tool calls and scoring rubrics, run against Claude or GPT-4, get scored reports via CLI and FastAPI dashboard. Lightweight, local-first, no lock-in.
A FastAPI WebSocket service and CI-friendly CLI for running LangSmith evaluations on demand — bundles LLM-as-judge (Claude) and heuristic evaluators, syncs datasets idempotently, and gates deploys via threshold checks.
Self-hostable LLM evaluation framework for measuring model performance across configurable skills. Run YAML-defined benchmarks against OpenAI/Anthropic models, score with LLM-as-judge, compare results in a CLI and web dashboard.
Autonomous Code Agent Tooling, Strict Eval Frameworks & Local-First CLI Utilities
🚀 基于Java的开源AI自动化评测框架 / An open source AI automation evaluation framework based on Java
Self-hosted evaluation framework for MCP servers. Define YAML test suites, run agents against MCP tools, score with LLM-as-judge rubrics, and monitor results in a FastAPI dashboard with tool-call traces and regression detection.
To associate your repository with the eval-framework topic, visit your repo's landing page and select "manage topics."