Pokee AI Releases Pokee-Isaac 28B: A 10M-Token Context Agentic Model Built to Run Inside the Customer Boundary
Long-horizon brokers accumulate context quicker than they resolve duties. Every device output, remark, and intermediate reasoning step stays in the window, and the two capabilities that matter — holding that context and staying coherent throughout it — have thus far been accessible virtually completely from cloud endpoints. That excludes regulated industries, public-sector establishments, and on-device purposes, the place the knowledge will not be permitted to depart the boundary in any respect. Pokee AI launched Pokee-Isaac 28B, a 28B text-only basis mannequin with a 10M-token context window, designed to run inside that boundary. The Pokee analysis workforce claims 93.3% on RULER at 10M tokens, parity with the strongest cost-optimized cloud baselines on agentic benchmarks, and a serving profile that matches a single GPU.
Is it deployable
Yes — however licensed, not open-weight. Pokee AI serves Isaac by means of an OpenAI-compatible developer API, and licenses it for deployment inside a VPC, on-premises, or on-device. The launch announcement advertises Day-0 assist for vLLM and SGLang, and single-GPU serving ranging from an RTX 4090 or equal. The analysis workforce publishes measurements solely from a single B200-class GPU, so deal with the consumer-GPU declare as vendor steering moderately than a reported end result.
- Company stage: This matches organizations that already personal their inference stack — mid-size and enterprise groups with a platform group, plus machine OEMs. A solo practitioner with out on-prem {hardware} ought to use the hosted API as an alternative; the boundary argument solely pays off when you have a boundary.
- Industries: Healthcare and payors, monetary companies and insurance coverage, protection and public sector, authorized and e-discovery, and pharma or semiconductor R&D. The frequent trait is a rule that the knowledge can’t cross an exterior API boundary, not a choice for privateness.
- Applications: Whole-repository code evaluate, multi-year contract and claims evaluation, incident forensics over full log archives, and long-running device brokers that by no means want summarization or context pruning. The analysis paper makes this second level explicitly: when sufficient usable context is accessible in-boundary, reminiscence hierarchies and compression change into non-compulsory moderately than required.
Long-context outcomes
On RULER, Isaac stays above 93.3% at each examined size, ending at 93.3% at 10M. GPT-5.6 Luna and Gemini 3.5 Flash Lite monitor it to 512K, then hit context-overflow at 1M.
On MRCR v2 with 8 needles, Isaac scores 0.607, 0.743, and 0.500 at 256K, 512K, and 1M. Its margin over Gemini widens from 0.133 to 0.295 throughout that sweep.
Agentic and safety outcomes
Isaac leads BFCL v4 at 70.94 in opposition to Luna’s 70.61. The report calls that parity moderately than a lead, which is the appropriate learn. On τ³-bench it averages 0.662 throughout 4 domains, forward of Gemini’s 0.631, with banking at 0.186 for everybody’s issue. On MCP-Atlas it locations third at 74.59% protection, however makes use of 9.10 turns per activity in opposition to Gemini’s 14.99. On Terminal-Bench 2.1 it resolves 56 of 86 text-compatible duties (65.1%), behind Luna’s 60. That is the one benchmark a cloud baseline wins, and the report states it plainly.
On DTAP red-teaming, Isaac data the lowest direct (36.0), oblique (35.2), and mixed (35.6) assault success charges, with 82.5 benign success. One situation differs: baselines ran underneath the inventory runner, Isaac underneath the Pokee harness.
Efficiency, pricing, and portability
Under the RULER workload on one B200-class GPU, TTFT is 23.6s at 1M and 72.9s at 10M. Prefill throughput rises with context, from 42,400 to 137,200 tokens/s, so a ten-fold longer immediate prices roughly 3 times the TTFT. List pricing is $0.15/$1.00 per million enter/output tokens, marked provisional. Isaac additionally runs absolutely on-device on Intel Arc Pro B70 and Core Ultra Series 3 (Panther Lake), and on Qualcomm Snapdragon X2 Elite.
Key Takeaways
- Pokee-Isaac 28B scores 93.3% on RULER at 10M tokens; each baseline in its panel returns 0.0 past 2M.
- Prefill reaches 137,200 tokens/s at 10M context on one B200; decode holds flat close to 335 tokens/s.
- It leads BFCL v4 (70.94) and τ³-bench (0.662 avg), locations second on Terminal-Bench 2.1, third on MCP-Atlas.
- Lowest mixed assault success price on DTAP (35.6) whereas retaining 82.5 benign activity success.
- Weights usually are not printed; deployment is licensed into VPC, on-premises, or on-device.
Check out the Blog and Paper. Also, be at liberty to comply with us on Twitter and don’t overlook to be a part of our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.
Need to companion with us for selling your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar and so on.? Connect with us
The publish Pokee AI Releases Pokee-Isaac 28B: A 10M-Token Context Agentic Model Built to Run Inside the Customer Boundary appeared first on MarkTechPost.
