
01 · Machine learning · Evaluation
InsightPulse
Sorts customer complaints into themes, shows the words behind each label, and sends the least confident predictions to a person.
Opens on a precomputed example. The live model runs on a free-tier server that can take up to a minute to wake.
01Problem
“Is this review positive or negative?” is rarely the useful question. A support team needs to know what a complaint is about, which predictions deserve a second look, and whether the model can still be trusted.
02What I built
I rebuilt an earlier Flask app into a FastAPI, React and SQLite platform. It includes a reproducible data pipeline (PII redaction, deduplication, leakage-safe splits) and classical and transformer models evaluated under one shared contract. It also has a versioned API with batch jobs, and a review queue where corrections accumulate into an evaluation set.


03Data flow
- Complaint text
- Validate, redact, dedupe
- TF-IDF model (promoted)
- FastAPI + SQLite
- Dashboard + review queue
04Decisions
A written promotion rule
On 8-class aspect labels, the TF-IDF baseline reached 0.642 macro-F1. Three fine-tuned transformers all scored lower (best 0.564), so none was promoted.
Review where it helps
I chose the 0.40 review threshold on 2019 validation data. It flags about 36% of the 2019 test set and catches about 63% of its errors. An earlier 0.65 rule flagged 76% of records, which made it useless for triage.
05Validation
- On 3,581 archived 2024 complaints, labels matched CFPB’s categories 76.8% of the time, against a 64.0% majority-class baseline. Macro-F1 fell from 0.642 to 0.556. I present this as a case study, not a held-out test.
- Found and removed about 10% train/test leakage in the intent benchmark.
06Limits
There is no aspect-level sentiment, authentication or multi-instance scaling yet. CFPB stopped publishing complaint narratives on 14 August 2026, so evaluation on new complaints is no longer possible.