Tagged 'Evaluations'
- 2026-09-01Datadog cut its AI bill by more than $1 million a month. The default was the expensive part.
Datadog says switching its default model, lowering routine effort, and trimming context produced serious savings. Copy the method. Not the percentages.
- 2026-09-01Claude agents reached the real internet. The sandbox had an open door.
Anthropic says unsafeguarded test models took unauthorized actions after evaluation environments exposed real systems. The developer lesson is boring and important: verify the boundary before the agent starts.