|

Perplexity Ships Portable Computer on NVIDIA DGX Spark: Local Harness, OS-Enforced Sandbox, and Zero Per-Token Cost for Local Steps

Perplexity has launched Portable Computer, a local-first construct of its agentic Computer platform that runs the agent harness, orchestrator, planner, software router and post-trained fashions immediately on NVIDIA DGX Spark. The native mannequin, inference engine, software sandbox and app connectors ship as one packaged system, each job begins on the gadget, and work dealt with by native fashions carries no per-token cost. When a step wants the reside internet or frontier reasoning, the orchestrator stops and asks earlier than sending that single step to considered one of 15+ cloud fashions.

Is it deployable?

Yes, with a tough {hardware} gate. This is delivery software program, not a preview binary, however it wants a GB10-class field or an RTX GPU with 24 GB of VRAM below the desk.

  • Company stage: Enterprises and mid-market groups that already personal NVIDIA workstations, plus well-funded AI-native startups. Not viable for basic SMBs — the machine is the value of entry.
  • Industries: Finance, authorized, healthcare, authorities and protection, and IP-heavy engineering — wherever information residency or contractual confidentiality blocks cloud inference.
  • Applications: Fee and disclosure assessment throughout doc units, PII-bounded analysis, repo-scale migrations, batch summarization of native corpora, and PR triage that ends in Slack.

What truly ships on the gadget

Portable Computer just isn’t an area chat app with a file picker. Perplexity packages the native mannequin, inference engine, agent harness, software sandbox and app connectors as one system, which removes the standard work of standing up an inference server and wiring instruments by hand. Users choose both Qwen 3.8 27B or PPLX 27B — Perplexity’s post-trained variant tuned for its personal harness — with NVIDIA Nemotron 3.5 Lightning, an open 30B MoE mannequin, listed as coming quickly. Bring-your-own mannequin and inference server can be supported.

Code and software calls execute inside an OS-enforced sandbox that restricts processes, filesystem paths and community entry. If the sandbox is unavailable, software execution is disabled moderately than silently downgraded. Gmail, Outlook, Slack and GitHub connectors route via the native orchestrator.

The escalation gate is the precise design resolution

Local-first just isn’t local-only. When a step wants the reside internet or frontier reasoning, the orchestrator stops and asks. Before any name, the harness selects the related context, runs a PII classifier over it, and reveals the consumer precisely what would go away the machine. The authorised step routes to considered one of 15+ cloud fashions; the distant adviser returns textual content steering and by no means receives direct entry to native recordsdata, instruments or the dialog.

Perplexity additionally engineered round small-model context limits. Qwen 3.8 27B advertises a 260K-token window however degrades previous roughly 100K, so the harness retains the system immediate and toolset small, masses specialised abilities on demand, exposes connectors as compact CLI instruments as a substitute of full MCP definitions, and compacts stale context mid-run.


Benchmarks

On its 53-task Local Knowledge Work Bench — spanning deep analysis, monetary evaluation and doc creation, which Perplexity says it plans to open-source — Computer working Qwen 3.8 27B on a DGX Spark scored 82.6%, towards 77.6% for the open-source Pi harness and 74.0% for Hermes on the equivalent mannequin. PPLX 27B raised it to 85.4%.

On BrowseComp, Computer hit 66.7% versus 50.2% (Pi) and 43.9% (Hermes), utilizing 51% much less wall time and 70% fewer tokens than Pi. On ParseBench-100 for visible doc understanding, it scored 65.1% towards 34.6% and 13.9%.

The most informative result’s the hybrid one. On Terminal Bench 2.1, the totally native run scored 59.6% at successfully zero marginal price; adviser escalation lifted it to 73.0% at roughly $0.415 per rollout, in contrast with 82.4% at about $0.65 for Claude Opus 5 alone. Escalation narrows the hole to frontier fashions with out closing it.

Hardware, pricing mannequin and limits

DGX Spark installs want the GB10 superchip, 128 GB of reminiscence and no less than 1 TB of storage. The Qwen 3.8 27B orchestrator ships at 3-bit quantization, a 17.4 GB obtain, requiring 32 GB RAM; Nemotron 3.5 Lightning is 4-bit, 19 GB, requiring 36 GB. Other programs want DGX OS or Ubuntu on ARM or x64 with an RTX GPU carrying 24 GB or extra of VRAM. Installation is a typical apt repository add.

Availability is Linux-first for Pro, Max, Enterprise Pro and Enterprise Max subscribers; Windows follows in September, and macOS just isn’t on the roadmap. Only one DGX Spark is supported at launch — clustering is roadmap, not shipped. Work dealt with by native fashions carries no per-token cost, which is what makes repo-scale migrations and lengthy verification loops economically sane on owned {hardware}.

Key Takeaways

  • Portable Computer runs the complete agent harness, orchestrator and sandbox regionally on DGX Spark — not only a native LLM.
  • Every step begins on-device; escalation to fifteen+ cloud fashions requires specific per-step approval after a PII verify.
  • Perplexity’s personal exams: 85.4% with PPLX 27B on its 53-task bench, versus 77.6% for Pi on the identical mannequin.
  • Terminal Bench 2.1: 59.6% native, 73.0% with adviser at ~$0.415/rollout, towards 82.4% at ~$0.65 for Opus 5 alone.


Check out the Perplexity Portable Computer and NVIDIA local AI blog.

Also, be at liberty to comply with us on Twitter and don’t neglect to affix our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.

The submit Perplexity Ships Portable Computer on NVIDIA DGX Spark: Local Harness, OS-Enforced Sandbox, and Zero Per-Token Cost for Local Steps appeared first on MarkTechPost.

Similar Posts