AI product evaluation and go or no go decisions
AI product decisions from pilot to go or no go
Five AI assisted workflows evaluated against real user problems, data quality, accuracy, adoption, and operational risk. Some capabilities were deployed, some remained pilots, and one was deliberately held back after testing.
Metrics are presented as approximate values from pilot and proof of concept testing. Customer and vendor details have been excluded.
- Role
- Senior Product Manager — problem framing, evaluation criteria, pilot design, and go or no go decisions
- Timeline
- 2023 to 2025
- Team
- Engineering, data science, analysts, and subject matter experts
Where AI belonged
Several slow workflows looked automatable. They did not carry the same risk.
Analysts spent time finding internal knowledge, interpreting complex analytical output, routing work to the right person, and asking questions of structured data. Each looked like a candidate for AI support.
The product question was never whether a model could produce an answer. It was where AI could assist safely, what needed human review, and which capabilities were not ready to move into production.
Knowledge retrieval assistant
Pilot
Problem
Users manually searched protocols, standard operating procedures, and project documentation.
Product approach
Retrieval augmented generation using approved SharePoint sources, with citations, prompt rules, and subject matter expert review.
- 5 to 6 pilot users
- Approximately 50 to 90 weekly queries
- Knowledge sources expanded from approximately 10 to 100 documents
A pilot with a small, observable user group. Not an enterprise wide deployment.
Analyst routing
POC
Problem
Project assignment depended on manual review of availability, experience, and therapeutic expertise.
Product approach
Tested routing recommendations using availability, past performance, and domain experience, with leadership retaining final approval.
Outcome
Estimated assignment time reduction of approximately 60 percent during pilot testing.
The 60 percent figure is an estimate from proof of concept testing, not a production metric.
AI summarization
Limited deployment
Problem
Users needed help interpreting complex analytical outputs.
Product approach
Generated structured summaries connected to dashboard and analytical results.
Evidence
In supported interpretation workflows, turnaround went from approximately 30 to 45 minutes to 15 to 20 minutes.
Limitation
Adoption and long term usage were not consistently tracked.
ICD code recommendation
Deployed with human review
Product approach
Recommended relevant codes and compared suggestions against a controlled reference dataset.
Limitation
Performance was weaker for rare diseases and ambiguous clinical terminology.
Text to SQL
POC, not advanced to production
Problem
Business users wanted natural language access to structured data.
Testing result
The POC generated inaccurate outputs and hallucinated schema references.
Decision
The feature was not moved into production without stronger schema controls, validation, access boundaries, and evaluation.
Product principles
How I decide whether AI ships.
- 01Start with the user decision, not the model.
- 02Define evaluation criteria before the pilot.
- 03Keep human review where errors have meaningful consequences.
- 04Treat a no go decision as a valid product outcome.
- 05Do not confuse technical capability with production readiness.