Writer launched a new flagship model called Palmyra X6 on Thursday, built as a post-trained variation of Z.ai's open source GLM-5.2, and paired it with an upgraded agentic harness the company says can cut customer costs by as much as 50% on basic tasks. The pitch is aimed squarely at enterprise buyers who are watching their token bills climb faster than any benchmark score justifies. Both the model and the harness upgrades are available to Writer clients starting Thursday.
Writer sells AI tools and agents to marketers, and CEO May Habib framed the release as a direct response to customer fatigue with the frontier-model treadmill. The company argues that most enterprise workloads don't need the largest available model — they need predictable spend on multi-step tasks that get executed reliably and with fewer tokens.
The strategic bet is that harness optimization, not model choice, is where the real cost savings live. A recent Writer research paper tested small changes in harness efficiency across multiple models and found the tweaks reduced costs by an average of 40% across the sample. The claim is that this compounds: whatever model an organization runs today or swaps to next quarter, the harness improvements travel with it.
Key facts
- 01Writer launched Palmyra X6 on Thursday, a post-trained variation of Z.ai's open source GLM-5.2 model aimed at enterprise deployments.
- 02The company estimates Palmyra X6 plus its upgraded agentic harness can cut customer costs by up to 50% on basic tasks.
- 03A Writer research paper found harness tweaks alone reduced costs by an average of 40% across the models it tested.
- 04Palmyra X6 runs alongside Writer's other models and third-party systems imported through Azure or Amazon Bedrock.
- 05CEO May Habib says CIOs are 'giving up on the labs' amid what she calls an unprecedented cost explosion.
Palmyra X6 itself is positioned as a deployment-ready model rather than a benchmark chaser. Building on GLM-5.2 lets Writer inherit the per-token economics of an open source base while layering its own post-training for enterprise tasks. Combined with the harness gains, Writer projects a 50% cost reduction for basic workloads compared to what customers are paying today.
For Writer's clients, the setup stays model-agnostic. Palmyra X6 sits alongside Writer's other in-house models and third-party systems that customers import through Azure or Amazon Bedrock, meaning teams don't have to rip out existing pipelines to test the new stack. That flexibility is part of the sales argument: buyers can route the cheap, fast tasks to Palmyra X6 and keep frontier models in reserve for the workloads that actually require them.
“The harness is the one component whose efficiency multiplies across every model an organization runs—present and future.”— Writer researchers, Writer research team
The underlying market signal is more interesting than the individual product. Enterprises deploying AI at scale are running into token bills that scale linearly with usage, and every additional agentic step multiplies the count. A harness that trims 40% off that multiplier changes the unit economics of running agents in production — which is exactly the workload category Writer is trying to own.
Habib's read on the frontier labs is unusually blunt. She argues that model providers have a structural incentive to drive up token consumption, and that chief information officers are noticing. The complaint isn't just about price — it's about whether the labs understand how enterprises actually deploy their models.
“The cost explosion here is just unprecedented for customers, and so is the degree to which CIOs are giving up on the labs.”— May Habib, CEO of Writer
That framing puts Writer in a specific competitive lane. Rather than out-benchmark OpenAI, Anthropic, or Google, the company is pitching itself as the layer that makes any of them — plus open source alternatives — cheaper to run. It's a wedge that only works if the harness gains hold up across real customer workloads, not just internal testing. If they do, the 40% harness figure is the more durable number in the announcement than the 50% headline.
The open question is how much of the cost savings actually reach customer invoices versus getting absorbed into Writer's margins as it scales. Enterprises that have been burned by opaque token pricing will want to see the reductions on their own bills before they treat the 50% number as a commitment. Habib's argument is that Writer needs to deliver visible savings or lose to the next vendor making the same promise.
The Palmyra X6 release lands in a market where cost, not capability, is quickly becoming the buying criterion for enterprise AI. If Writer's harness thesis is right — that the orchestration layer, not the model, is where the durable savings sit — then the frontier labs face a squeeze from both sides: open source models eroding per-token pricing at the bottom, and infrastructure vendors like Writer capturing the efficiency gains at the top. That's a harder position to defend than a benchmark lead.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




