Skip to main content
Live
Main content

Quilty bets AI can predict box-office hits — and gets Sinners wrong

The startup's script-scoring tool rated flop Christy above Sinners, which grossed $370 million and won an Oscar.

Jaeden Schafer
Editor in Chief · · 5 min read
Quilty bets AI can predict box-office hits — and gets Sinners wrong

Quilty, an AI startup pitching Hollywood on script-level box-office prediction, rated the script for Christy — which grossed around $2 million — higher than the script for Sinners, which grossed $370 million and went on to win an Oscar. The miss, reported by The Verge after sessions with the founders, undercuts the company's central claim that its tool can read an unproduced script and forecast commercial success. Quilty charges $50 per analysis and returns a 0-to-100 score covering narrative quality, commercial viability, audience resonance, and projected production budget.

Founded by film producers Simon Horsman and Daniel Wood, Quilty does not train its own model. It routes scripts through a stack of off-the-shelf systems: Gemini for structural breakdowns, a US-hosted instance of DeepSeek for financial modeling, and a combination of Claude and ChatGPT for narrative and character analysis. A few minutes after a writer uploads a script, the platform spits out a report with story beats, character notes, and an estimated budget.

The pitch to writers, producers, financiers, and studio executives is that a Quilty score can serve as an objective signal in the greenlight process — high scores as a calling card, low scores as a flag for revisions. The founders frame this as democratizing access for up-and-coming creatives rather than replacing development executives.

Key facts

  • 01Quilty charges $50 per script analysis and returns a 0–100 score covering narrative, commercial viability, audience resonance, and budget.
  • 02The tool rated Christy, which grossed roughly $2 million, higher than Sinners, which grossed $370 million and won an Oscar.
  • 03Quilty does not train its own model — it routes scripts through Gemini, ChatGPT, Claude, and a US-hosted DeepSeek instance.
  • 04Founders Simon Horsman and Daniel Wood blamed the Sinners miss on Sydney Sweeney's star power weighting in Christy's favor.
  • 05The platform uses VADER, an open-source sentiment library, as part of what the founders call its 'sentiment engine.'

Horsman told The Verge the company is trying to keep humans in the loop rather than automate development outright. He said the founders solicited feedback from creatives wary of generative AI's effect on jobs while building the product.

We agree with a lot of the negative sentiment towards AI, but what we're trying to do is enable human creativity.
Simon Horsman, Quilty co-founder

The architecture is modular by design. Wood, who serves as CTO, said the idea grew out of his own experience using consumer chatbots to handle a real estate lawsuit a few years back, after ChatGPT refused to help him and told him to go find a lawyer. He moved to Gemini for its longer context window, then tested Grok after seeing Elon Musk post about its legal benchmark scores. Each model, he concluded, was good at different things.

That logic carries through to Quilty's stack. Rather than fine-tuning a single model, the company picks the best general-purpose system for each subtask and uses context prompting to suppress hallucinations. Wood argues this makes the platform upgradeable on someone else's R&D budget.

He also signaled the company has no loyalty to US frontier labs if Chinese models start winning on benchmarks. That position is unusual among Hollywood-facing vendors, most of which avoid associating their tooling with non-US model providers.

When Claude Mythos comes out and I can see that it's a better LLM, all of a sudden, my whole software gets better.
Daniel Wood, Quilty co-founder and CTO

The harder problem is whether the underlying premise works. Quilty's founders concede the platform missed Sinners because, on paper, Sweeney's draw plus the lower production cost of a boxing biopic looked like the safer financial bet than a fantasy-action feature. That is a defensible piece of pre-release logic and also exactly the kind of conventional-wisdom call human development executives have been making — and getting wrong — for a century.

Related · from this week
Spotify's AI DJ adds French, German, Italian and Brazilian Portuguese
Jaeden Schafer · 4 min read →

The founders also acknowledged blind spots Quilty cannot model. The platform could not have anticipated that Magazine Dreams, which Horsman produced, would be derailed by the legal troubles of its lead actor Jonathan Majors in 2023. Nor could it have predicted the Chicken Jockey meme cycle that drove A Minecraft Movie's outsized run. The sentiment engine, which leans on the open-source VADER library to score text valence, has no read on what will become a phenomenon on TikTok between the script lock and opening weekend.

Skeptics will note that script-based prediction is one of the oldest and most-attempted problems in entertainment analytics, and that no firm — AI-driven or otherwise — has solved it. Large language models are pattern recognizers trained on existing text; asking them to forecast the reception of art that does not yet exist is a category mismatch. A high Quilty score on a script that flops, or a low score on one that breaks out, is not a tunable error — it is the product working as designed on a problem it cannot answer.

The market opportunity is real even if the product is not yet credible. Studios spend millions on coverage, focus groups, and tracking, and a $50 per-script tool that produces a polished report is genuinely cheaper than the human equivalent. If Quilty's scores improve at the pace of the underlying frontier models — which is the founders' explicit bet — the gap between marketing claim and actual predictive power will narrow. For now, the Christy-versus-Sinners result is the headline the company has to live with, and it is the kind of public miss that makes financiers slower, not faster, to trust the number on the report.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Tools

Spotify's AI DJ adds French, German, Italian and Brazilian Portuguese
Tools

Spotify's AI DJ adds French, German, Italian and Brazilian Portuguese

The interactive feature jumps from 2 to 6 languages and lands in 8 new markets, reaching more than 75 countries.

Jaeden Schafer4 min read
Google logo
Security

Gemini, ChatGPT and Claude are surfacing real phone numbers, privacy firm says

DeleteMe says AI-related privacy requests jumped 400% in seven months, with ChatGPT cited in 55% of complaints and Gemini in 20%.

Jaeden Schafer5 min read
Amazon's search bar now generates AI images of products you can't buy
Tools

Amazon's search bar now generates AI images of products you can't buy

The in-app feature surfaces invented clothing and home goods as you describe them, letting shoppers tap an image to find real lookalikes.

Jaeden Schafer4 min read