AI Box released a Model Context Protocol server this week that plugs its catalog of 80+ AI models directly into Claude, ChatGPT, Gemini, and Cursor. The server exposes 60 flagship endpoints across text, image, video, and audio — including OpenAI's Sora 2, Google's Veo 3.1, Black Forest Labs' FLUX.2, and ElevenLabs — callable by name from whichever assistant a user already prefers. Plans start at $7.25 a month.
The launch closes a specific gap for Claude, which has no native image, video, or audio generation. Through the MCP connection, Claude can hand off mid-conversation to gpt-image-1.5, Sora 2 Pro, Veo 3.1 Fast, or ElevenLabs Music, then keep reasoning over the returned assets. The traffic runs both ways: ChatGPT users can call Claude Opus 4.8 through AI Box, and Gemini users can call GPT-5.6.
The pitch is that model loyalty stops being a lock-in. Every provider ships its own chat app and refuses to expose competitor models inside it; MCP routes around that by letting a third-party server publish tools the host assistant can invoke.
“Every one of these companies wants you to live inside their app, and none of them will ever give you their competitor's best model.”— Jaeden Schafer, Founder of AI Box
Key facts
- 01AI Box's MCP server exposes 80+ models with 60 flagship endpoints callable by name across text, image, video, and audio.
- 02The server ships 11 tools, including four direct-generation tools and a plain-English workflow builder called Vibe Builder.
- 03The roster spans 19 text models, 24 image models, 8 video models, and 9 audio models from providers including OpenAI, Anthropic, Google, xAI, and ElevenLabs.
- 04Pricing starts at $7.25/month for Starter (annual), with Standard at $16, Pro at $32, and Business at $64.
- 05Generated media returns as presigned URLs valid for 24 hours, with async polling for long video jobs and a spend gate on high-cost runs.
The server ships 11 tools. Four are direct generation — chat completion, image, video, and audio — each with a model picker. A fifth, called Vibe Builder, takes a plain-English workflow description and compiles it into a reusable "box" with a typed input schema, so a pipeline like "blog post in, thumbnail plus 30-second promo plus voiceover out" can be re-run with new inputs indefinitely. The remaining tools cover box search, box execution by ID, run history across MCP and scheduled jobs, in-flight polling, and one-call replay of any previous run.
The direct-call roster covers 19 text models from seven providers, including OpenAI's GPT-5 line through GPT-5.6 with Luna, Sol, and Terra variants; Anthropic's Claude Fable 5, Opus 4.8, Sonnet 5, Sonnet 4.6, and Haiku 4.5; Gemini 3 Pro and Flash; xAI's Grok 4.5 and 4.3; Meta's Llama 3.3 70B; DeepSeek V3; and Mistral Large. The image list runs to 24 models across the FLUX family, gpt-image-1.5, Gemini 3 Pro Image, ByteDance's Seedream variants, Qwen-Image, Ideogram, and Stable Diffusion XL. Video adds 8 models — Sora 2 and Sora 2 Pro, Veo 3.1 and Veo 3.1 Fast, Seedance 1.0 Pro and Lite, PixVerse V5, and Grok Imagine Video. Audio adds 9, spanning ElevenLabs' TTS, music, sound effects, speech-to-speech, and speech-to-text, plus OpenAI TTS, Whisper transcription and translation, and Grok TTS.
The practical workflows are the point. A writer can draft in Claude and generate a hero image with FLUX.2, an Ideogram thumbnail, and a 15-second Sora 2 teaser in the same session. Podcasters can produce cover art, ElevenLabs intro music, sound effects, and Whisper transcripts without leaving chat. Developers in Cursor can generate app icons and UI mock imagery mid-session; assets return as URLs ready to paste into a repo. Marketers can build one box — say, a product-launch kit with 10 catalog shots, a launch video, and a jingle — and rerun it per SKU with reproducible seeds.
“The thing people miss about MCP is that it's not just chat — it's workflows.”— Jaeden Schafer, Founder of AI Box
The engineering choices matter for real use. Long video jobs run asynchronously: the tool returns a run ID and the assistant polls until the file lands, so the chat stays responsive. Finished media comes back as presigned URLs valid for 24 hours, refreshable from run history. A spend gate requires explicit confirmation before any generation whose estimated cost crosses a user-set threshold — an agentic loop can't quietly drain the budget. Idempotency keys prevent double-billing on flaky connections, and users can set a default model per modality so casual prompts need zero configuration.
Founder Jaeden Schafer framed the launch as breaking a captive-app dynamic. "You keep the assistant you love and get everybody's models behind it — Claude users can generate Sora videos now. That was science fiction six months ago," he said. On the workflow tooling, he added: "You describe a pipeline once in plain English, AI Box turns it into a box, and from then on your assistant can run it, schedule it, and replay it. That's the difference between a toy and a tool you build a business on."
The obvious constraint is that AI Box now sits between the user and every underlying provider — meaning outages, rate limits, or pricing changes at OpenAI, Anthropic, or Google flow through a single vendor. The 24-hour URL expiry also means production pipelines need to persist assets rather than rely on the returned links, and the spend gate is only as useful as the threshold the user sets. Enterprises evaluating this will want to look hard at data-handling terms with 13-plus upstream providers running through one contract.
For the broader market, aggregators are the sleeper category of the MCP era. Frontier labs will keep refusing to surface competitors inside their own chat surfaces because captive usage is the moat; MCP lets a third party rebuild the multi-model experience users actually want, and charge a flat subscription for it. If enough of that value accrues to aggregators — one login, one bill, every model — the labs eventually face a choice between opening their own front ends or watching the assistant relationship migrate to whoever routes the models.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




