Thursday, 24 Sep 2026
Subscribe to AIWatcher
AIWatcher
  • Home
  • News

    OpenAI slots two cheaper GPT-6 models beneath Astra

    By
    AIWadmin

    Researcher hijacks Meta’s Muse agent through a hidden setting

    By
    AIWadmin

    ChatGPT learns to take voice orders for office chores

    By
    AIWadmin

    YouTube hands creators an AI agent for titles and thumbnails

    By
    AIWadmin

    Listeners can now rewrite the Spotify algorithm in plain words

    By
    AIWadmin

    Anthropic’s Claude agents flag a new enzyme family in viral DNA

    By
    AIWadmin
  • Articles

    Data shop Snorkel AI banks $350M as labs stockpile training sets

    By
    AIWadmin

    Ema banks $77M to push AI employees into the back office

    By
    AIWadmin

    US military command randomises routes to outfox enemy models

    By
    AIWadmin

    Kyutai teaches a speech model to do arithmetic out loud

    By
    AIWadmin

    Nokia hands developers a no-training route to calibrated answers

    By
    AIWadmin

    An open model sorts eight overlapping speakers in real time

    By
    AIWadmin
  • Spotlight

    DeepMind keeps cloud memory encrypted behind device-held keys

    By
    AIWadmin

    Anthropic trims Opus costs and speeds output in a mid-cycle refresh

    By
    AIWadmin

    Cisco Talos builds a fingerprint library for AI-driven malware

    By
    AIWadmin

    Google parks idle agents in a new open source runtime

    By
    AIWadmin

    OpenAI gives mathematicians a voice but hands them no brake

    By
    AIWadmin

    NVIDIA teaches its robotics stack to take instructions from agents

    By
    AIWadmin
  • About
    • Mission
    • Services
    • Contact
  • Newsletter
  • Shop
    • All Items
    • By Category
      • Hats
      • T-Shirts
    • Cart
  • 🔥
  • Alignment
  • Classification
  • Distillation
  • Explainability
  • Hallucination
  • Legal/Compliance
  • Medical
  • NLM
  • Mobility
  • Research
  • Robotics
  • Safety
  • Startups
  • Prompt
  • Python
  • RAG
  • RLHF
  • Token
  • Vision
Font ResizerAa
AIWatcherAIWatcher
  • Home
  • News
  • Articles
  • Spotlight
  • About
  • Newsletter
  • Shop
Search
  • Home
  • News
  • Articles
  • Spotlight
  • About
    • Mission
    • Services
    • Contact
  • Newsletter
  • Shop
    • All Items
    • By Category
    • Cart
Have an existing account? Sign In
Follow US
© 2022 Foxiz News Network. Ruby Design Company. All Rights Reserved.
Articles

Google trains a diffusion retriever to widen a single search

Google Research has distilled reinforcement-learned retrieval behaviour into a small diffusion model that fans queries out 12 to 20 times faster.

AIWadmin
Last updated: September 18, 2026 1:17 am
AIWadmin
ByAIWadmin
Global AI news & information.
Follow:
Share
SHARE

Ask a system to widen one vague query and it may hand back the same idea wearing two different outfits.

Google Research calls that paraphrastic collapse, and it is one of two problems the team identifies with letting a general-purpose model generate query fan-out at inference time. The example in the paper starts from bohemian festival style and produces bohemian festival fashion and festival bohemian clothes, near-duplicates that retrieve almost identical results. The second problem is speed: autoregressive generation plus repeated retrieval calls is slow, and sampling several candidates multiplies the cost.

Their answer, Retrieve-for-Train, learns good fan-out once with reinforcement learning and then distils the behaviour into a small diffusion model that produces every retrieval direction in a single pass.

Rewards differ by task. For open-ended abstract retrieval, three weighted terms combine, and an ablation showed all three are load-bearing: groundedness alone drove the policy into repetitive strings, alignment on its own accelerated collapse as the model echoed the query back, and only diversity closed both shortcuts. Training used GRPO with soft PPO regularisation adding forward and reverse KL penalties, a group size of 8 and a global batch of 512.

Quality was judged by a model on a 5-point scale. On Polyvore, the Gemma3-4B variant averaged 49.1 against 40.9 for best-of-N and 38.5 zero-shot, while diversity climbed from 56.0 to 76.8. The diffusion version held 74.3.

The timing difference is starker still. At batch size 8, autoregressive fan-out needed about 1.46 seconds against 0.07; at 1,024 the gap stretched to nearly 50 seconds against 4.21, a 12x to 20x advantage.

TAGGED:Diffusion ModelsGoogle ResearchReinforcement LearningResearchretrievalSearch
SOURCES:MarkTechPost
Share This Article
Email Copy Link Print
ByAIWadmin
Follow:
Global AI news & information.
Previous Article Stanford turns published papers into runnable research agents
Next Article MIND banks $72M to chase data leaks inside AI tools

You Might Also Like

Articles

Kyutai teaches a speech model to do arithmetic out loud

By
AIWadmin
News

Google gives publishers a button to collect loyal readers

By
AIWadmin
Articles

Pinecone open-sources a recipe book for vector quantization

By
AIWadmin
News

Harvey previews its first in-house legal model for long case runs

By
AIWadmin
AIWatcher
Facebook Twitter Youtube Linkedin Rss

Global AI News and Information
AIWatcher is your definitive source for AI updates worldwide, from Silicon Valley to Shanghai.
Our industry coverage keeps you in the loop with the latest news and trends shaping the future of AI.

Quick Links
  • News
  • Articles
  • Spotlight
  • Events
About Us
  • Mission
  • Services
  • Contact
  • Privacy Policy
  • Legal
© 2026 AIWatcher. All Rights Reserved.