Monday, 17 Aug 2026
Subscribe to AIWatcher
AIWatcher
  • Home
  • News

    Needle 2 squeezes tool calling into a 14MB model binary

    By
    AIWadmin

    Dyna-2 robot model learns from 170 years of human video

    By
    AIWadmin

    Kog bets software unlocks 30x faster inference on stock GPUs

    By
    AIWadmin

    Trust deficit, not dire warnings, drives AI backlash, Amodei says

    By
    AIWadmin

    Gas price tripling forecast clouds hyperscaler power plans

    By
    AIWadmin

    DeepSeek hikes V4 prices as Flash flubs complex agent jobs

    By
    AIWadmin
  • Articles

    New $300M fund backs AI for hospitality and venues

    By
    AIWadmin

    Ransomware operators put Claude Code inside live hacks

    By
    AIWadmin

    Claude Code bills for blank thinking blocks users never see

    By
    AIWadmin

    Liquid AI’s 3B vision model brings screen agents to laptops

    By
    AIWadmin

    Okta trims agent token bills by hiding unneeded MCP tools

    By
    AIWadmin

    Samsung trains wearable health models that learn from biosignals

    By
    AIWadmin
  • Spotlight

    Suno Studio 2.0 lets keyboards feed AI tracks with real playing

    By
    AIWadmin

    Google lets creators drop visible watermarks from Gemini output

    By
    AIWadmin

    Anthropic keeps Model 2 in the vault as misalignment risk ticks up

    By
    AIWadmin

    OpenAI crosses $40B run rate as enterprise outgrows consumer sales

    By
    AIWadmin

    Databricks tops $7B run rate as $5B round values it at $190B

    By
    AIWadmin

    SpaceX closes $60B Cursor deal to pair code agents with GPU fleet

    By
    AIWadmin
  • About
    • Mission
    • Services
    • Contact
  • Newsletter
  • Shop
    • All Items
    • By Category
      • Hats
      • T-Shirts
    • Cart
  • 🔥
  • Alignment
  • Classification
  • Distillation
  • Explainability
  • Hallucination
  • Legal/Compliance
  • Medical
  • NLM
  • Mobility
  • Research
  • Robotics
  • Safety
  • Startups
  • Prompt
  • Python
  • RAG
  • RLHF
  • Token
  • Vision
Font ResizerAa
AIWatcherAIWatcher
  • Home
  • News
  • Articles
  • Spotlight
  • About
  • Newsletter
  • Shop
Search
  • Home
  • News
  • Articles
  • Spotlight
  • About
    • Mission
    • Services
    • Contact
  • Newsletter
  • Shop
    • All Items
    • By Category
    • Cart
Have an existing account? Sign In
Follow US
© 2022 Foxiz News Network. Ruby Design Company. All Rights Reserved.
News

NVIDIA’s Nemotron 3.5 Lightning routes agent grunt work to small models

NVIDIA released Nemotron 3.5 Lightning, a 30B open model for the busy work inside agent runs, plus a router that picks the best model per step.

AIWadmin
Last updated: August 12, 2026 7:31 pm
AIWadmin
ByAIWadmin
Global AI news & information.
Follow:
Share
SHARE

NVIDIA has released two open pieces of the agent stack: a small fast model for the busy work inside agent runs, and a router that picks the right model for each step.

Nemotron 3.5 Lightning packs 30B parameters but activates only 3B per token, using a Mamba-2, MoE, and attention hybrid with a one-million-token context window. On output speed the vendor claims a 4x edge over similarly sized models, and a 30 percent gain on 10,000 PinchBench tasks relative to Qwen3.6 35B at similar accuracy. A single GPU, from a DGX Spark to an H100, can serve it.

Lightning is aimed at the execution layer of long-running agents: tool calls, result validation, subagent delegation, and review routing, where routing every step to a frontier reasoning model wastes time and money. Firms across cybersecurity, legal, coding, finance, and healthcare are adapting it, including CrowdStrike, Harvey, CodeRabbit, Fastino Labs, and Lila Sciences. The model is out now under the permissive OpenMDW-1.1 license, with open weights, training data, and recipes shipped alongside.

NeMo Switchyard, its companion, routes each workflow step to the most capable and efficient model using tuning-free routers, including one built on an LLM classifier. NVIDIA says internal tests cut task costs to a third. Both tools anchor an August campaign pushing local AI, joined by the open-weights LTX-2.5 world model.

TAGGED:AI AgentsLocal AImodel routingnemotronNvidiaopen source
SOURCES:MarkTechPostVentureBeat
Share This Article
Email Copy Link Print
ByAIWadmin
Follow:
Global AI news & information.
Previous Article AI agents found a Zoom takeover bug in fewer than 20 prompts
Next Article Oracle clouds a 98-qubit quantum computer for AI workloads

You Might Also Like

News

AI Pioneer LeCun Warns of Industry Bubble, Calls Musk’s xAI a Misstep

By
AIWadmin
News

Microsoft and OpenAI Ditch the AGI Clause. The Hype Was Always the Point.

By
AIWadmin
News

Musk’s OpenAI Bid Reveals His Need for Absolute Control

By
AIWadmin
News

New $300M fund backs AI for hospitality and venues

By
AIWadmin
AIWatcher
Facebook Twitter Youtube Linkedin Rss

Global AI News and Information
AIWatcher is your definitive source for AI updates worldwide, from Silicon Valley to Shanghai.
Our industry coverage keeps you in the loop with the latest news and trends shaping the future of AI.

Quick Links
  • News
  • Articles
  • Spotlight
  • Events
About Us
  • Mission
  • Services
  • Contact
  • Privacy Policy
  • Legal
© 2026 AIWatcher. All Rights Reserved.