Context
A high-volume complaints and correspondence corpus that executives could describe anecdotally but not measure.
Constraint
Free-text quality varied by channel and decade, and any classification error carried reputational consequence.
Approach
A production NLP workflow with topic and entity analysis, human review of low-confidence classifications, and an executive insight layer with published method.
What made it hard, really
Accuracy on the long tail of rare topics stayed low. We published the confidence bands instead of averaging them away.
Risk controls
Privacy, security, data quality, model governance, resilience and human oversight were tested as acceptance criteria, with control owners named in the delivery plan and evidence retained for audit.
Reusable assets and IP
Reference patterns and reusable assets are described without exposing proprietary implementation. Client-specific code, documentation and controls remain client-owned.