A Google engineer presents techniques for deploying tiny LLMs (sub-billion parameter models) as on-device agents using the Google AI Edge stack. Key topics include the LiteRT/LiteLM runtime, a skill harness built on top of Gemma 4 via AI Core for agent-style function calling, and fine-tuning tiny models with synthetic data. A concrete example shows fine-tuning Function Gemma (270M parameters) to improve function-calling accuracy from 46% to over 90% for 8 out of 10 target functions. The talk also covers a production transcription app (Eloquent) built with chained tiny models, and the Google AI Edge Gallery open-source app for testing models on Android and iOS.
•21m watch time
19 Impressions