Together.ai has rolled out major updates to its Batch Inference API, including a streamlined UI for creating and tracking batch jobs, universal support for all serverless models and private deployments, a 3000× rate limit increase (from 10M to 30B enqueued tokens per model per user), and pricing at 50% of the real-time API cost. Key use cases include large-scale text analysis, synthetic data generation, embedding generation, fraud detection, content moderation, model evaluation, and customer support automation.

2m read timeFrom together.ai
Post cover image
Table of contents
What's New