Andrew Mercer
on this page

Overview

This guide targets Product Managers/Owners who both build AI-powered products and use AI tools to manage their own agile workflow — two related but distinct skill sets that are increasingly expected of the same role.

Who it's for: Product Owners/Managers working on or adjacent to AI-driven features, operating within a Scrum or Kanban team.

Core Topics Breakdown

1. Product Management for AI Features (Building AI Products)

  • Writing user stories for probabilistic/AI-driven features differs from deterministic features: acceptance criteria must account for a range of acceptable outputs, not a single expected result.
  • Defining "Definition of Done" for an ML/AI feature often needs to include model evaluation metrics (accuracy, precision/recall, or user-satisfaction thresholds) alongside standard functional criteria.
  • Managing the unique risk profile of AI features: hallucination/error rates, bias, and the need for human review loops — these should appear explicitly in the backlog as non-functional requirements, not afterthoughts.

2. Backlog Structuring for AI/ML Work

  • Distinguishing between exploratory spikes (data exploration, model feasibility) and committed delivery stories — spikes are time-boxed investigations, not guaranteed-outcome commitments, and should be planned accordingly.
  • Sequencing: data pipeline and labeling work often needs to precede feature stories that depend on model output quality.
  • Communicating uncertainty to stakeholders — AI feature timelines are often less predictable than standard CRUD feature work, and the Product Owner needs a vocabulary for expressing this without losing stakeholder confidence.

3. Using AI Tools to Manage the Agile Process Itself

  • Drafting user stories and acceptance criteria with AI as a first pass, refined by the Product Owner and team together.
  • Summarizing customer feedback (support tickets, reviews, survey responses) at scale using AI to spot backlog-worthy themes faster than manual review.
  • Generating competitor/market research summaries to inform prioritization — always followed by human verification of the underlying facts and sources.

4. Metrics for AI-Driven Products

  • Beyond standard product metrics (adoption, retention), track model-specific health metrics: response latency, error/hallucination rate, user override/correction rate (how often users reject the AI's suggestion).
  • Feedback loops: designing a mechanism for users to flag bad AI output, and feeding that data back into the backlog as prioritized bug/improvement work.

5. Ethical and Governance Considerations

  • Bias auditing as a recurring backlog item, not a one-time launch gate.
  • Transparency to users about when they're interacting with an AI feature (a recurring theme in emerging regulation across jurisdictions).
  • Data privacy considerations specific to training or fine-tuning on user data.

Study Tips

  • Practice rewriting a standard user story ("As a user, I want X so that Y") into an AI-feature-aware version that includes acceptable-output-range criteria and an evaluation metric.
  • Build a mental checklist for AI feature Definition of Done: functional criteria + evaluation metric threshold + human-review/fallback path + bias/privacy check.
  • Study at least one real public example of an AI feature failure (e.g., a chatbot giving harmful advice) and identify what backlog/process gap likely contributed.

Common Pitfalls

  • Writing AI feature stories with the same rigid acceptance criteria used for deterministic features, which sets the team up to "fail" against an unrealistic standard.
  • Treating model evaluation as a one-time pre-launch task instead of an ongoing backlog item as real-world data drifts over time.
  • Under-communicating uncertainty to stakeholders, leading to broken trust when AI feature timelines slip.

Quick Reference Cheat Sheet

Concept Definition
Spike Time-boxed investigation, not a committed deliverable
Model drift Real-world data diverging from training data over time, degrading accuracy
Human-in-the-loop A review/override mechanism for AI output

Further Practice

Rewrite a standard feature user story as an AI-feature story, including an evaluation metric and a fallback/human-review path in its Definition of Done.