GenAI evaluation · LLM-as-a-judge
Quality- and cost-aware GenAI evaluation for Signify Health
For Signify Health, a CVS Health company, I built Azure ML foundations, a Snowflake call-center rating application, and an LLM-as-a-judge architecture designed around judge quality per dollar.
- +8% predictive accuracy. Signify Health, a CVS Health company.
- +15% conversion rate. Signify Health, a CVS Health company.
- 25% faster call resolution. Signify Health, a CVS Health company.
Context & problem
Signify Health, a CVS Health company, needed a consistent way to rate call-center interactions while balancing predictive quality, serving cost, and operational usefulness. I delivered the work through Tredence.
What I built
Three layers, in sequence:
- Platform foundation: Set up the Azure ML platform supporting the ML work.
- The application: Built a Snowflake-based LLM application rating call-center interactions in production.
- The evaluation layer: Designed an LLM-as-a-judge architecture on Snowflake with serving economics as a first-class constraint, optimizing judge quality per dollar.
Measured result
+8% predictive accuracy, +15% conversion rate, and 25% faster call resolution versus the prior scoring approach.
Reach out to learn more.