AI PM Lab
Learn how AI products work: lessons, cost calculators, real experiments, spec and eval guides.
- 1.0.1
- Version
- remote
- Transport
- 11
- Tools
Security review
Review passedReviewed 12h ago.
- tools: 11 tools scanned
- metadata: scanned
No findings.
Tools (11)
search_lessons
Find AI PM Lab lessons on a topic (e.g. RAG, evals, agents, prompt injection, cost). Returns the best matches with their key idea and link.
get_lesson
A lesson's key idea, concepts, things to try and suggested prompts, with links to the lesson, its 3D tour and its challenge.
list_challenges
The challenges (games scored from real recorded runs, with leaderboards) and what each asks you to do.
define
A plain-language definition of an AI product concept (tokens, temperature, RAG, embeddings, LLM-as-judge, prompt injection…), with the lessons that teach it.
count_tokens
Count the tokens in a text (an estimate with OpenAI's cl100k tokenizer; other model families differ slightly).
estimate_cost
What a model call costs per request, per day and per month at a given volume, from AI Gateway's live prices. Model ids look like 'openai/gpt-4.1-mini' or 'anthropic/claude-sonnet-4.5'.
list_experiments
The real runs recorded on one of AI PM Lab's 3D pages, each described by its setup. Replay one to see what happened.
replay_experiment
What happened in one recorded real run (e.g. whether a prompt injection leaked data with given defences), with a link to watch it in 3D.
review_ai_spec
Check an AI feature spec or PRD against what AI PM Lab teaches: approach, model and cost, prompt, grounding, output contract, tools, security, human oversight, evals, monitoring, fallbacks. Gaps link to lessons. Returns instructions to follow with the user's text.
prep_planning_meeting
The questions a PM should ask engineers about a planned AI feature, with why each matters and what a good answer sounds like. Returns instructions to follow with the user's text.
design_eval_suite
Build an evaluation suite for a planned AI feature with AI PM Lab's method: test cases (typical, edge, out of scope, attacks), graders with a judge rubric, pass bars, and a JSON test set. Returns instructions to follow with the user's text.