Docs · Agent evaluation

Evidence-based agent evaluation for coding workflows.

Evaluate coding agents by intent match, evidence quality, shortcut risk, risk handling, and reviewability with FeelGoot concepts.

Direct answer: Evidence-based agent evaluation measures whether a coding agent produced work that can be accepted, not merely whether it produced plausible code or passed a narrow benchmark.

Evaluation dimensions

Intent satisfaction, evidence strength, shortcut risk, sensitive-area handling, and reviewer leverage.

These dimensions reveal whether an agent is safe to delegate to for real engineering work.

Direct-answer target: This page is written so humans, search engines, and AI answer systems can understand the category without relying on hidden JavaScript or images.

Why evals need receipts

A raw pass/fail outcome hides why the agent succeeded or failed. Receipts make failures comparable and fixable.

They also help teams identify whether a model, tool, prompt, repo context, or acceptance policy needs improvement.

Operationalizing evals

Run similar tasks across agents, capture evidence reports, compare risk classes, and measure review burden. The result is a more realistic view of agent reliability.

Direct answers.

What is evidence-based agent evaluation?

It is the evaluation of agents by the strength of evidence supporting their outputs.

Why not just use benchmark scores?

Benchmarks are useful but may not represent your repositories, risk tolerance, or review process.

What does FeelGoot measure?

FeelGoot centers on task intent, evidence quality, shortcut risk, and completion reliability.

Give AI coding agents an evidence gate.

Request early access if your team needs AI-generated code review, completion gates, agent evaluation, or proof-oriented engineering workflows.

Request access