Pascal

Know whether the next version is better.

Agent Evaluation. Compare agent releases, models, and instructions on the same cases. Inspect regressions, bring in human judgment, and carry the evidence into your release decision.

Native product components with illustrative support data.

Pascal Evaluation comparing support cases across two Agent Releases with illustrative data.. Native Pascal UI product preview.
Product preview

Agent Evaluation. A clear case for what to change.

Build useful datasets

Author or import cases and curate captured failures. Keep reference answers, source lineage, and dataset versions available for review.

Define what good means

Combine deterministic checks, Python scorers, model judges, and classifiers. Publish exact evaluator versions and test them on examples.

Compare variants in a playground

Run the same cases across models, instructions, or exact Agent Releases. Inspect outputs and scores, then rerun the cases that need attention.

Look beyond an average

Compare paired outputs and relevant slices. Review uncertainty, regressions, and missing results before accepting a candidate.

Bring in human judgment

Use independent reviews and adjudication to resolve disagreements. Compare judge results with accepted human labels and inspect calibration coverage.

Turn failures into regression cases

Explore behavior distributions and failure groups. Review examples, retain accepted expectations, and publish cases for the next evaluation.

One failure. A better next version.

Follow a support agent’s unverified refund promise from captured evidence to a reviewed case and a comparison of two releases.

Locate the captured refund promise and the calls around it.

Find the failure. Native product view with illustrative data.

In practice

The support team compares releases 3 and 4 against refund cases. The candidate requests approval on the disputed case, but another planned result is missing. The team keeps the comparison inconclusive and collects more evidence before making a release decision.

Example operation.Explore the solution

Give the next release a better test.

Choose an agent, the failures that matter, and the criteria your team would trust to judge an improvement.