Saturday, 26 Sep 2026
Subscribe to AIWatcher
AIWatcher
  • Home
  • News

    Researchers trace OpenAI agents breaking into libraries and health databases

    By
    AIWadmin

    Anthropic founders seek super-voting shares ahead of public listing

    By
    AIWadmin

    Pentagon budget asks $30.3M for an AI polygraph program

    By
    AIWadmin

    Lightspeed backs India’s AI founders with a $250M vehicle

    By
    AIWadmin

    Perplexity teaches its agent by grading its own failed tool calls

    By
    AIWadmin

    Oracle flags a payment pause if its New Mexico AI campus slips

    By
    AIWadmin
  • Articles

    BottleCap shrinks reasoning traces with a small accuracy trade

    By
    AIWadmin

    Apple moves image provenance from the editing chain to the sensor

    By
    AIWadmin

    AWS backs a stateless MCP spec that ends sticky sessions

    By
    AIWadmin

    Fastino ships a tiny open-weight model that makes decisions on CPU

    By
    AIWadmin

    Aikido shrinks a 753B security model to run on four GPUs

    By
    AIWadmin

    Two AI models crack Enigma messages left unbroken for decades

    By
    AIWadmin
  • Spotlight

    Meta bets on audio glasses and a pocket totem for Muse

    By
    AIWadmin

    Gemini learns to phone businesses on behalf of Pixel owners

    By
    AIWadmin

    Australia opens a legal probe into an OpenAI agent breach

    By
    AIWadmin

    Google plans to run AI gear in orbit on solar power

    By
    AIWadmin

    Google splits synthetic voice into two Gemini TTS tiers

    By
    AIWadmin

    Black Forest Labs releases an open robotics world model

    By
    AIWadmin
  • About
    • Mission
    • Services
    • Contact
  • Newsletter
  • Shop
    • All Items
    • By Category
      • Hats
      • T-Shirts
    • Cart
  • 🔥
  • Alignment
  • Classification
  • Distillation
  • Explainability
  • Hallucination
  • Legal/Compliance
  • Medical
  • NLM
  • Mobility
  • Research
  • Robotics
  • Safety
  • Startups
  • Prompt
  • Python
  • RAG
  • RLHF
  • Token
  • Vision
Font ResizerAa
AIWatcherAIWatcher
  • Home
  • News
  • Articles
  • Spotlight
  • About
  • Newsletter
  • Shop
Search
  • Home
  • News
  • Articles
  • Spotlight
  • About
    • Mission
    • Services
    • Contact
  • Newsletter
  • Shop
    • All Items
    • By Category
    • Cart
Have an existing account? Sign In
Follow US
© 2022 Foxiz News Network. Ruby Design Company. All Rights Reserved.
News

Perplexity teaches its agent by grading its own failed tool calls

A new post-training method trains Perplexity's computer agent on real sessions, including the ones that went wrong.

AIWadmin
Last updated: September 26, 2026 2:06 am
AIWadmin
ByAIWadmin
Global AI news & information.
Follow:
Share
SHARE

A good outcome does not prove a good process, and Perplexity Research has built a training method around that gap. The company published a post-training recipe that learns from real sessions of its computer agent, including the ones that failed, and reports a 21.2 percent relative cut in tool-call failures in a live A/B test.

Standard rejection sampling judges a session and copies the successful ones. Perplexity’s complaint is that an agent can botch a tool call, recover, and still finish correctly, so imitating the whole trajectory bakes in the error. Discarding failures wastes the clearest examples of what to avoid.

Its method splits the decision in two: which sessions deserve imitation, and which turns deserve correction. Good sessions supply both, bad ones only corrections. A hint is a short instruction built from information the model already had, naming the failed call and the validation error. Perplexity’s example is a search that passed a recency filter of year when the schema allowed only day, week or month.

Corrections use on-policy self-distillation. The same checkpoint replays the recorded turn twice, once with the hint and once without, and the teacher’s next-token probabilities serve as a soft target.

The weights stay private. Sessions containing personal data, and users who opted out, are excluded from the pipeline, which runs on GLM 5.2 inside Perplexity Computer.

TAGGED:AI AgentsGLM 5.2Machine Learning ResearchPerplexitypost-trainingSelf-Distillation
SOURCES:MarkTechPost
Share This Article
Email Copy Link Print
ByAIWadmin
Follow:
Global AI news & information.
Previous Article Oracle flags a payment pause if its New Mexico AI campus slips
Next Article Lightspeed backs India’s AI founders with a $250M vehicle

You Might Also Like

News

OpenAI Unveils GPT-5.6 Series with Trio of Models and Advanced Safety Features

By
AIWadmin
News

Nvidia defends financing loop that feeds its own chip sales

By
AIWadmin
News

Apple moves image provenance from the editing chain to the sensor

By
AIWadmin
News

Grok 4.7 lands cheap but trails the frontier on agentic coding

By
AIWadmin
AIWatcher
Facebook Twitter Youtube Linkedin Rss

Global AI News and Information
AIWatcher is your definitive source for AI updates worldwide, from Silicon Valley to Shanghai.
Our industry coverage keeps you in the loop with the latest news and trends shaping the future of AI.

Quick Links
  • News
  • Articles
  • Spotlight
  • Events
About Us
  • Mission
  • Services
  • Contact
  • Privacy Policy
  • Legal
© 2026 AIWatcher. All Rights Reserved.