A new VB Pulse survey of 157 enterprise AI teams reveals a widening evaluation gap: AI agents are being granted increasing autonomy while the testing systems meant to verify them remain widely distrusted. Half of respondents report that an agent or LLM feature passed internal evaluations but still caused a customer-facing failure — one in four experienced this more than once.
Despite these failures, 66% of organizations already permit some production deployment without human review or plan to do so within 12 months. Only 5% say they fully trust the automated evaluations that would justify those release decisions. The mismatch has created what researchers call a “control gap” — organizations shipping agentic systems faster than they can build safety layers around them.
The most common reason cited for distrusting automated evaluation is poor alignment with real-world outcomes, flagged by 29% of respondents. Bias and inconsistency followed at 21%, with lack of explainability at 18%.
The findings suggest the next year will be a retrofit cycle, with enterprise buyers shifting budget toward governance, identity, and evaluation systems that make agentic deployments manageable. VentureBeat explored the topic further at its Transform 2026 conference in Menlo Park.