All work

AI product evaluation and go or no go decisions

AI product decisions from pilot to go or no go

Five AI assisted workflows evaluated against real user problems, data quality, accuracy, adoption, and operational risk. Some capabilities were deployed, some remained pilots, and one was deliberately held back after testing.

Metrics are presented as approximate values from pilot and proof of concept testing. Customer and vendor details have been excluded.

Role
Senior Product Manager — problem framing, evaluation criteria, pilot design, and go or no go decisions
Timeline
2023 to 2025
Team
Engineering, data science, analysts, and subject matter experts
01

Where AI belonged

Several slow workflows looked automatable. They did not carry the same risk.

Analysts spent time finding internal knowledge, interpreting complex analytical output, routing work to the right person, and asking questions of structured data. Each looked like a candidate for AI support.

The product question was never whether a model could produce an answer. It was where AI could assist safely, what needed human review, and which capabilities were not ready to move into production.

02

Knowledge retrieval assistant

Pilot

Problem

Users manually searched protocols, standard operating procedures, and project documentation.

Product approach

Retrieval augmented generation using approved SharePoint sources, with citations, prompt rules, and subject matter expert review.

  • 5 to 6 pilot users
  • Approximately 50 to 90 weekly queries
  • Knowledge sources expanded from approximately 10 to 100 documents

A pilot with a small, observable user group. Not an enterprise wide deployment.

03

Analyst routing

POC

Problem

Project assignment depended on manual review of availability, experience, and therapeutic expertise.

Product approach

Tested routing recommendations using availability, past performance, and domain experience, with leadership retaining final approval.

Outcome

Estimated assignment time reduction of approximately 60 percent during pilot testing.

The 60 percent figure is an estimate from proof of concept testing, not a production metric.

04

AI summarization

Limited deployment

Problem

Users needed help interpreting complex analytical outputs.

Product approach

Generated structured summaries connected to dashboard and analytical results.

Evidence

In supported interpretation workflows, turnaround went from approximately 30 to 45 minutes to 15 to 20 minutes.

Limitation

Adoption and long term usage were not consistently tracked.

05

ICD code recommendation

Deployed with human review

Product approach

Recommended relevant codes and compared suggestions against a controlled reference dataset.

Limitation

Performance was weaker for rare diseases and ambiguous clinical terminology.

06

Text to SQL

POC, not advanced to production

Problem

Business users wanted natural language access to structured data.

Testing result

The POC generated inaccurate outputs and hallucinated schema references.

Decision

The feature was not moved into production without stronger schema controls, validation, access boundaries, and evaluation.

07

Product principles

How I decide whether AI ships.

  1. 01Start with the user decision, not the model.
  2. 02Define evaluation criteria before the pilot.
  3. 03Keep human review where errors have meaningful consequences.
  4. 04Treat a no go decision as a valid product outcome.
  5. 05Do not confuse technical capability with production readiness.