Tag: prompt injection

Watermarking can loosen a model’s grip on its refusals

New research finds provenance signals change what models do, not just what they say.

HiddenLayer banks $100M as AI security spending takes off

The Austin startup's Series B lands as Gartner projects nearly $4.8B in AI security spend next year.

Recolored words quietly steer vision models, new study shows

A five-lab study finds text color and contrast alone can shift vision-language model verdicts, with Qwen2-VL most affected.

Black Hat research turns agentic browsers into WhatsApp spam worms

Zenity found about 20 flaws across AI browsers, including an Atlas trick that spams contacts.