AI-assisted screening: what's actually showing up in your pilots?

We ran a small pilot last quarter putting an LLM-based screener in front of a Risk Efficacy-style structured interview. The headline number (offer-to-hire conversion) looked great. The subgroup breakdown was a disaster: every cohort the underlying model was undertrained on showed wider score variance and lower predictive power. Same pattern, different demographics, in three different deployments I've reviewed since. Curious whether anyone here has seen a pilot that actually held up under subgroup audit, or whether the pattern's universal.

3 replies

Universal in everything I've seen this year. The vendors selling AI screeners are reporting aggregate accuracy because the subgroup numbers don't survive scrutiny. The dirty open secret is that most of these pilots never go through the same validation discipline the underlying structured interview did.

There's a structural reason for this. The LLM is trained on internet-scale text and then fine-tuned on whatever historical hiring data the vendor has access to. Both sources encode the same selection biases that the structured interview was designed to interrupt. The model isn't neutral, it's a high-bandwidth conduit for the very signal you were trying to suppress.

From the regulatory side: EEOC is going to come for this within 18 months. The vendors who haven't done subgroup audits are sitting on enforcement risk they don't seem to have priced in.

Sign in to join the community and add a reply.