Engineering leaders from Workato, Hippocratic AI, and ISMG shared production lessons from running high-volume AI inference workloads at DigitalOcean Deploy 2026. Key themes include: tool selection accuracy degrades sharply when AI agents have access to 50+ tools; P99 latency becomes a patient-safety issue in clinical voice applications; AI agents should never have admin-level permissions but instead operate as time-scoped per-action delegates; and companies that delay structuring their data and workflows before adopting AI risk falling two years behind on their operating model. The panel's consensus is that scaling inference is fundamentally an infrastructure and governance problem, not a model problem.
Table of contents
The Built for Mass Scale PanelistsAI Has Gone From “Secret Sauce” to Standard InfrastructureWhat Works at Ten Requests Fails at a MillionThe Agentic Identity CrisisThe Latency TrapPlan for the Architecture That Hasn’t Shipped YetTighten Agent Permissions to Shrink the Blast RadiusThe Riskiest AI Strategy Is No AI Strategy283 Impressions