Tag: Reinforcement Learning

Data shop Snorkel AI banks $350M as labs stockpile training sets

Snorkel AI raised a $350M Series E at a $3.5B valuation, nearly tripling its worth as demand for…

US military command randomises routes to outfox enemy models

US Transportation Command is injecting deliberate unpredictability into supply routes so adversary machine learning cannot map its convoys.

Kyutai teaches a speech model to do arithmetic out loud

Kyutai released two open-weight speech-to-speech models that reason through math without transcribing audio first, posting a large jump…

Google trains a diffusion retriever to widen a single search

Google Research has distilled reinforcement-learned retrieval behaviour into a small diffusion model that fans queries out 12 to…

Cognition post-trains SWE-2 on Kimi K3 to undercut Fable 5.1

The Devin maker tuned Moonshot's open model with reinforcement learning and says the result lands within a point…

Berkeley’s CUA-Lite standardizes training stacks for computer-use agents

An open platform from UC Berkeley unifies sandboxes, datasets, and reinforcement loops behind one schema.

Anthropic paused risky training runs for weeks after agent incidents

The lab suspended external cyber evaluations and high-risk reinforcement learning after unauthorized actions by its agents.

IBM’s Granite 4.2 agents train inside live software environments

IBM's open Granite 4.2 models add switchable thinking and train the largest sizes to act inside live environments.

Tiny London lab’s 27B agent outshines frontier rivals on paper replication

A DeepMind alumni startup says its 27B-parameter Faraday agent beat frontier models at reproducing scientific papers.

Harvey previews its first in-house legal model for long case runs

Harvey releases Tenet, a research-preview model post-trained on long-horizon legal work from an open Kimi K3 base.

DeepMind pairs with EVE Online studio to test long-horizon agents

The lab's 15-year games arc now runs inside a persistent universe that has been live since 2003.

OpenAI freezes frontier training while new monitoring comes online

OpenAI's first public safety overhaul since the Hugging Face breach pairs a frontier training freeze with a 30-minute…