Evaluate a prompt against a dataset
Turn real user feedback into a dataset, run an experiment that sweeps prompt versions and models, and read the run report.
Improve a prompt from feedback
Turn replies your team disagreed with into candidate rewrites of your production prompt, score them against what is live, and promote the one that earns it.
Evaluate with conversation history
Carry a session's prior turns into a dataset example so a run, the judge, and the optimizer see the same conversation the model actually answered in.
Score live traffic
Create a standing rule that judges matching production gateway calls automatically, with no dataset and no manual run.