Infrastructure & Governance
AI-assisted screening: what's actually showing up in your pilots?
May 10, 2026
We ran a small pilot last quarter putting an LLM-based screener in front of a Risk Efficacy-style structured interview. The headline number (offer-to-hire conversion) looked great. The subgroup breakdown was a disaster: every cohort the underlying model was undertrained on showed wider score variance and lower predictive power. Same pattern, different demographics, in three different deployments I've reviewed since. Curious whether anyone here has seen a pilot that actually held up under subgroup audit, or whether the pattern's universal.