Friction

Feature request · Integrations · Blocks work

Support decision models as AI Eval judges

1 source thread · first seen 2026-09

Summary

AI Evals only supports chat-model judges, making bounded evaluations costly and noisy for the author’s use case. They request decision-model support with bounded questions and probability results, or at minimum validation and an error for unsupported models.

Affects
Teams using AI Evals with OpenRouter
Workaround
Run the decision model in their own backend and send the result to PostHog as a custom event.

Evidence

Excerpts are copied word for word from the source; follow the link to read it in full.

“Selecting `typesafe/jev-1.13` as the OpenRouter model of an LLM-judge evaluation today produces no `$ai_evaluation` result”

Report this item