A hands-on demonstration of running large language models locally on Apple M5 Max (128GB RAM) using LM Studio. The video compares a 120B parameter GPT-OSS model against the newer Qwen 3.6 35B model for agentic browser automation tasks using the Playwright MCP server. The GPT-OSS model struggles significantly with tool calling, taking ~5 minutes for simple browser interactions, while Qwen 3.6 completes the same login-and-employee-creation workflow rapidly with zero tool call failures, correctly selecting the right tool for each action. Key highlights include MLX format acceleration on Apple Silicon, Qwen 3.6's think preservation feature, and its strong agentic coding capabilities.

15m watch time