DigitalOcean has launched Evaluations, a native LLM-as-a-Judge evaluation feature integrated directly into its Inference Engine. Teams can validate models, fine-tunes, BYOM imports, and inference router configurations against their own datasets before pushing to production. The feature includes six pre-built metrics (correctness, completeness, faithfulness, PII, toxicity, bias), custom rubrics, reusable evaluation presets, MCP support for programmatic triggering in CI/CD pipelines, and versioned dataset management supporting CSV and JSONL up to 1GB. Billing is token-based; dataset and result storage is free for the first 12 months. Premium commercial models (OpenAI, Anthropic) require a tier 2 account.
Table of contents
DigitalOcean Evaluations CapabilitiesHow to Access EvaluationsStart Evaluating Before You Ship621 Impressions