Monday, 17 Aug 2026
Subscribe to AIWatcher
AIWatcher
  • Home
  • News

    Needle 2 squeezes tool calling into a 14MB model binary

    By
    AIWadmin

    Dyna-2 robot model learns from 170 years of human video

    By
    AIWadmin

    Kog bets software unlocks 30x faster inference on stock GPUs

    By
    AIWadmin

    Trust deficit, not dire warnings, drives AI backlash, Amodei says

    By
    AIWadmin

    Gas price tripling forecast clouds hyperscaler power plans

    By
    AIWadmin

    DeepSeek hikes V4 prices as Flash flubs complex agent jobs

    By
    AIWadmin
  • Articles

    New $300M fund backs AI for hospitality and venues

    By
    AIWadmin

    Ransomware operators put Claude Code inside live hacks

    By
    AIWadmin

    Claude Code bills for blank thinking blocks users never see

    By
    AIWadmin

    Liquid AI’s 3B vision model brings screen agents to laptops

    By
    AIWadmin

    Okta trims agent token bills by hiding unneeded MCP tools

    By
    AIWadmin

    Samsung trains wearable health models that learn from biosignals

    By
    AIWadmin
  • Spotlight

    Suno Studio 2.0 lets keyboards feed AI tracks with real playing

    By
    AIWadmin

    Google lets creators drop visible watermarks from Gemini output

    By
    AIWadmin

    Anthropic keeps Model 2 in the vault as misalignment risk ticks up

    By
    AIWadmin

    OpenAI crosses $40B run rate as enterprise outgrows consumer sales

    By
    AIWadmin

    Databricks tops $7B run rate as $5B round values it at $190B

    By
    AIWadmin

    SpaceX closes $60B Cursor deal to pair code agents with GPU fleet

    By
    AIWadmin
  • About
    • Mission
    • Services
    • Contact
  • Newsletter
  • Shop
    • All Items
    • By Category
      • Hats
      • T-Shirts
    • Cart
  • 🔥
  • Alignment
  • Classification
  • Distillation
  • Explainability
  • Hallucination
  • Legal/Compliance
  • Medical
  • NLM
  • Mobility
  • Research
  • Robotics
  • Safety
  • Startups
  • Prompt
  • Python
  • RAG
  • RLHF
  • Token
  • Vision
Font ResizerAa
AIWatcherAIWatcher
  • Home
  • News
  • Articles
  • Spotlight
  • About
  • Newsletter
  • Shop
Search
  • Home
  • News
  • Articles
  • Spotlight
  • About
    • Mission
    • Services
    • Contact
  • Newsletter
  • Shop
    • All Items
    • By Category
    • Cart
Have an existing account? Sign In
Follow US
© 2022 Foxiz News Network. Ruby Design Company. All Rights Reserved.
News

Pokee’s single-GPU model opens a ten-million-token window

Regulated industries get cloud-class long context without sending data anywhere.

AIWadmin
Last updated: August 9, 2026 10:48 pm
AIWadmin
ByAIWadmin
Global AI news & information.
Follow:
Share
SHARE

Long-context AI has largely been a cloud game, and that has left regulated industries on the sidelines. Pokee AI thinks it has the answer with Pokee-Isaac 28B, a licensed text model whose 10-million-token window is designed to run inside a customer’s own VPC, on-premises hardware, or devices, with no data crossing an external API boundary.

The company reports 93.3 percent on the RULER benchmark at the full 10 million tokens and claims parity with the strongest cost-optimized cloud baselines on agentic evaluations. Serving needs a single GPU, with day-zero support for vLLM and SGLang, and Pokee says an RTX 4090-class card can get started, though its published measurements come from one B200. The model ingests longer prompts faster, not slower: aggregate prefill speeds rise from about 42,400 tokens per second at 1 million tokens of context to 137,200 at 10 million, while decode stays near 335 tokens per second. On its agentic scorecard Isaac’s 70.94 sits barely ahead of GPT-5.6 Luna’s 70.61, a lead the technical report describes as a tie.

Healthcare payors, financial services, defense, legal, and pharma R&D are the target buyers, and the use cases run long: full-repository code review, multi-year contract analysis, and incident forensics over entire log archives. Pokee’s argument is that once the full context lives inside the boundary, teams no longer need memory hierarchies or compression schemes to cope. The trade-off is the license itself, since Isaac is not open-weight.

TAGGED:Agentic AIAI InfrastructureEnterprise AIlong contextmodel releaseon-prem AIPokee-Isaac
SOURCES:MarkTechPost
Share This Article
Email Copy Link Print
ByAIWadmin
Follow:
Global AI news & information.
Previous Article NVIDIA-backed Firebird flips on a 70,000-GPU campus in Armenia
Next Article Legal AI startup Harvey nears a $15.5B valuation

You Might Also Like

News

Meta’s Brain2Qwerty v2 Translates Thoughts into Text Without Surgery

By
AIWadmin
News

Mira Murati’s Testimony Reveals Deep Rot at OpenAI’s Core

By
AIWadmin
News

AI Boom Reshapes Global Semiconductor Industry, Sparks Regional Divergence and In-House Chip Surge

By
Zoe Chang
News

Dictation AI Arms Race: Who Wins When Your Keyboard Becomes Obsolete?

By
AIWadmin
AIWatcher
Facebook Twitter Youtube Linkedin Rss

Global AI News and Information
AIWatcher is your definitive source for AI updates worldwide, from Silicon Valley to Shanghai.
Our industry coverage keeps you in the loop with the latest news and trends shaping the future of AI.

Quick Links
  • News
  • Articles
  • Spotlight
  • Events
About Us
  • Mission
  • Services
  • Contact
  • Privacy Policy
  • Legal
© 2026 AIWatcher. All Rights Reserved.