# Pascal RL

Source: https://www.trypascal.io/platform/rl

Make the next response better informed.

Turn reviewed responses into evidence for improvement. Capture which answer was better, why it was preferred, and what the next candidate needs to demonstrate.

## Capabilities

### Choose the behavior to improve

Choose a recurring decision or failure. Define the result you want to improve, the examples to review, and how the judgments will be used.

### Keep the original comparison

Compare the original responses with their source context. Hide model identities and scores during review so people can judge the answers on their merits.

### Make the rubric explicit

Define the review criteria, rating scales, and questions. Keep the rubric version with each response so later readers can understand the judgment.

### Resolve disagreement

Collect independent judgments and apply the agreed review rules. Send unresolved disagreements to another reviewer.

### Retain the accepted case

Retain the final judgment, the reviews behind it, and the exact source example. Give future evaluations a case the team can inspect.

### Propose the next improvement

Propose a next step for Evaluation, Training, or Serving. Each product checks the inputs, permissions, and review needed before acting.

## Connected context

Bring response pairs from Models and Evaluation into review. Retain the original source and accepted judgment so later evaluation or training can use the case under its own requirements.

- Improvement objectives, source examples, and review rubrics
- Independent reviews, disagreements, and accepted judgments
- Evidence references and proposals for further evaluation or training

## Controls and permissions

Human review records a judgment. It does not train, deploy, or authorize a model change. Each later step needs its own inputs, evaluation, and approval. The walkthrough shows an accepted review, not a self-updating model.

## Example

Both responses see fresh telemetry, but only one respects the pending approval. Reviewers prefer the response that keeps the shipment held and retain the source and rubric behind that choice. A future improvement can start with an inspectable case instead of an unexplained thumbs-up.

## Workflow

### Compare the responses

- One response keeps the Compliance hold. The other would tender while approval is still pending. Current telemetry satisfies only one of the required conditions.

### Inspect the basis for the judgment

- Expand the retained source item and rubric version. Reviewers can see the evidence the preference was based on.

### Keep the accepted evidence

- The preference remains connected to its source and review. It is evidence for further improvement; this comparison does not train a model or publish a release.

## Next steps

Choose a recurring failure, define what a better response would do, and collect reviewed examples with the evidence needed to judge them.

- [Contact sales](https://www.trypascal.io/contact)
- [Pascal Work](https://www.trypascal.io/solutions/pascal-work)
