TeardownCadre STATE appliqué21 juillet 202610 min de lectureMorgan Stanley AI @ Work Assistant STATE Teardown: Evals Are Not TracesMorgan Stanley runs one of the most disciplined eval programs in enterprise AI: pre-deployment evals, human-graded tests, a daily regression suite. The STATE score is still 4/10, because nothing in the public record can reconstruct what the assistant told one advisor about one client yesterday. This teardown dissects the gap between aggregate evals and per-interaction traces, and shows the single table that closes it.
TeardownCadre STATE appliqué6 juillet 20269 min de lectureRamp Financial Automation Agent STATE Teardown: Graduated Autonomy Proves the Decision, Not the WorkflowRamp runs one of the most disciplined AI agent rollouts in production finance — graduated autonomy, human escalation, 99% policy accuracy. STATE score: 6/10. Across three public sources, there is no evidence the approval workflow survives a crash mid-flight.