Perplexity has brought its Portable Computer agent to Windows PCs running NVIDIA GeForce RTX and NVIDIA RTX PRO Workstations, letting users run multistep AI workflows locally on GPUs with 24GB or more of VRAM. The release extends existing support for NVIDIA DGX Spark systems and RTX PCs running Linux, and it lets users keep sensitive files off the cloud while avoiding consumption of Perplexity Computer credits. The default model on the local side is Qwen 3.8 27B, post-trained by Perplexity for the agent and tuned for NVIDIA RTX GPUs.
Portable Computer is the on-device sibling of Perplexity Computer, the cloud agent that plans and executes multistep tasks. On a supported RTX machine it can analyze files, cross-reference documents and handle recurring work without a chat transcript leaving the PC. When a task needs heavier reasoning than the local model can deliver, the agent flags it and asks the user for permission before routing anything to a cloud model.
“As local models become more capable, AI agents can handle more work directly on a PC while keeping sensitive information on the device.”— Gerardo Delgado, NVIDIA
The setup pitch is that a Windows user does not have to assemble the model stack themselves. Perplexity ships the local model, the runtime and the Computer tooling — including the built-in browser and the proprietary SPACE sandbox — as a single install through the Perplexity app. That is a meaningful concession to normal users, who have watched the local-model ecosystem fragment across Ollama, LM Studio, llama.cpp and a growing list of quantization formats.
Key facts
- 01Perplexity Portable Computer launched on Windows for NVIDIA GeForce RTX and RTX PRO GPUs with 24GB or more of VRAM.
- 02The agent runs a locally-optimized Qwen 3.8 27B model post-trained to work with Perplexity Computer.
- 03Local execution keeps sensitive files on-device and does not consume Perplexity Computer cloud credits.
- 04Connectors ship for Outlook, OneDrive, Word, Google Drive, Gmail, Slack and GitHub.
- 05Support for NVIDIA DGX Station is expected to follow the initial Windows release.
Connectors extend the agent into everyday workflows: Microsoft Outlook, OneDrive, Word, Google Drive, Gmail, Slack and GitHub are all wired in at launch. NVIDIA's example workloads span engineering (reviewing open pull requests in a GitHub project and flagging outdated documentation), finance (parsing two years of brokerage summaries, 1099s and tax returns with per-page citations) and product analytics (analyzing a funnel export to explain a drop in activation and posting the summary to Slack).
The privacy pitch is the clearest commercial argument. A financial workflow that walks through consolidated 1099s and tax returns is exactly the kind of task where enterprise buyers balk at sending documents to a hosted chatbot. Running it against a 27B-parameter model on the user's own workstation, with cloud escalation gated by explicit consent, removes that objection without forcing the user to give up the Computer agent's orchestration.
“Users can put the agent to work without having to research models or configure the complex software stack typically required to run local AI.”— Gerardo Delgado, NVIDIA
The launch sits inside a broader push by NVIDIA to make DGX Spark and RTX workstations the reference hardware for a new class of local agents. In parallel, Z.ai's GLM 5.3 Flash is being optimized for DGX Station and dual DGX Spark systems, and Z.ai's 744-billion-parameter GLM 5.3 flagship is tuned for hours-long agent sessions running on a DGX Station plus a cluster of four DGX Spark systems. Qwen has released Qwen 3.8-Flash-Next and an early preview of Qwen 4 that can run locally on a single DGX Spark using NVFP4 quantization. DeepSeek-v4.1 Flash is being pitched at agent workloads that need to cut key-value cache memory demands.
That lineup matters because it defines what an on-device agent stack can now realistically do. A year ago the answer was short summarization tasks with a 7B model. The current answer is a 27B agent that plans across a filesystem and a set of SaaS connectors, with a clean off-ramp to a frontier cloud model for the reasoning-heavy 10%. The gap between local and cloud is still real, but it is narrowing faster than most enterprise procurement cycles.
There are limits worth naming. The 24GB VRAM floor rules out most consumer laptops and a large share of prosumer desktops — this is a workstation-class release, not something that lights up every Copilot+ PC. DGX Station support has not yet shipped. And escalation to cloud models, while gated by consent, is still the path for anything that stretches Qwen 3.8 27B's reasoning ceiling, which means the fully-local story is bounded by what a mid-sized open model can actually do.
Perplexity's move sharpens a split in how agent platforms are being built. OpenAI, Anthropic and Google have leaned into cloud-hosted agents with browser control and hosted tool use. Perplexity is betting that a nontrivial share of high-value agent work — legal, financial, engineering, anything with regulated data — will migrate to the endpoint the moment the local model is good enough. Pairing that bet with NVIDIA's workstation channel gives it a distribution path that does not depend on the browser or the phone. The remaining question is whether the local model curve keeps closing the gap to frontier cloud reasoning fast enough to make on-device the default rather than the fallback.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.



