Morgan Stanley AI @ Work Assistant STATE Teardown: Evals Are Not Traces
Morgan Stanley runs one of the most disciplined eval programs in enterprise AI: pre-deployment evals, human-graded tests, a daily regression suite. The STATE score is still 4/10, because nothing in the public record can reconstruct what the assistant told one advisor about one client yesterday. This teardown dissects the gap between aggregate evals and per-interaction traces, and shows the single table that closes it.