1. Quality -- All five sections present and correctly structured, no section is a verbatim lift from the transcript (paraphrase, don't regurgitate), no hallucinated content — every claim traceable to something said in the session, appropriate length — long enough to be useful at review time, short enough that someone actually reads it. 2. Measuring success - percentage of users who open the summary x+ days after the session, e.g. target threshold to validate: If >40% of summaries generated are opened at least once after day 7, the feature is being used as intended. For the prompts, I prompt Claude to synthesise the actual prompts to use to generate the key learning summary and visual image first.