Tag: LLM Evaluation

AWS opens a benchmark that tests agents on live cloud work

AWS's new open-source benchmark runs agents against real resources in disposable cloud accounts.

DeepSeek hikes V4 prices as Flash flubs complex agent jobs

DeepSeek's V4 Flash tops leaderboards but passes barely half its agent tasks as API prices jump.

Arato Secures $10 Million in Seed Funding to Predict AI Failures Before They Happen

Arato platform simulates thousands of user interactions across text, voice, and image modalities to catch AI failures before…