A practical walkthrough of evaluating LLMs for media planning tasks using R's vitals package. The author, who works at Havas Media and created the HAI-Q benchmark, builds a small eval with questions, solvers (Claude, GPT), and LLM-as-judge scoring. The post explains why media-planning evals are uniquely hard: plans are underdetermined, media facts are market- and time-specific, terminology distinctions matter for calculations, and vitals alone doesn't cover agent trajectories. HAI-Q benchmark results show GPT-5 answers only 45.7% of 35 domain-specific questions correctly, Claude Sonnet 4.5 gets 17.1%, and GPT-4o just 11.4%. Numerical reasoning around reach and audience sizes is the biggest weakness. The takeaway: LLMs can accelerate media work but every numerical recommendation needs verification before use.