Evaluate assistant answers and agent actions continuously: grounded answers, correct abstentions, blocked restricted sources, stale-context invalidations, latency and model mix.
Visuals: Groundedness by route; Why answers were withheld or blocked; Model routes: quality vs latency; Agent actions awaiting approval