Skip to main content
Live
Main content

Sequent launches to fix alignment before superintelligence arrives

A new nonprofit from UK AISI and Timaeus researchers wants $100–150M to chase theoretical alignment guarantees the frontier labs aren't pursuing.

Jaeden Schafer
Editor in Chief · · 5 min read
Sequent launches to fix alignment before superintelligence arrives

A new alignment research nonprofit called Sequent launched this week with a blunt thesis: the empirical safety work at frontier AI labs will not deliver confidence that superintelligence is safe before it is built. Sequent is targeting $100–150M in initial funding and 40–80 full-time employees within a couple of years, with founders drawn from the UK AI Security Institute's alignment team and the theory startup Timaeus. The organization says it is prepared to raise at least one order of magnitude more if it can show progress across many parallel research bets.

The launch landed alongside Cognition's new FrontierCode coding benchmark and a Chinese cultural-reasoning evaluation called ChinaHeritaQA, both surfaced in this week's Import AI by Jack Clark. Together they sketch a research week defined by the gap between rising model capability and the alignment work meant to keep it in check.

Artificial superintelligence (ASI) may be developed in the next few years. It is unclear whether alignment is on track to be ready on the same timeframe. At a minimum, the empirical programs at AI labs are unlikely to deliver a priori confidence, before training ASI, that things will go well.
Sequent, alignment research nonprofit

Sequent's pitch is structural. The group argues that frontier-lab safety work is, in its words, essentially reactive, producing methods that function but offer no principled insight into when they will fail. Sequent wants something stronger: principled reasons for being confident that the alignment we observe in situations we control generalizes to alignment in situations we cannot easily control, such as large-scale, long-horizon tasks executed in the world. The research portfolio spans scalable oversight, learning theory, heuristic arguments, game theory, and personas, with the bet that interactions between those threads will produce results none of them generate alone.

Key facts

  • 01Sequent is targeting $100–150M in initial funding and 40–80 full-time employees within a couple of years.
  • 02The nonprofit pulls staff from the UK AI Security Institute alignment team and theory startup Timaeus.
  • 03Cognition's new FrontierCode benchmark tops out at 13.4% for Claude Opus 4.8 on its 50-task Diamond tier.
  • 04ChinaHeritaQA, a new VLM benchmark with 14,133 QA pairs across 51 UNESCO sites, sees Qwen-VL-8B-Instruct hit 81% versus 67% for humans.
  • 05Xiaomi's MiMo-V2.5-Pro-UltraSpeed is a 1 trillion parameter LLM serving 1,000 tokens per second.

The funding ask is modest next to the frontier labs but unusually ambitious for an independent safety nonprofit. Sequent's stated goal of $100–150M initially, with a 10x follow-on if the work pans out, puts it in the same conversation as Anthropic's early safety-focused rounds, though without commercial revenue to backstop it. The team frames independence as the point: an outside organization can raise the alarm if frontier labs ship something dangerous. Sequent's own line is that, when needed, we might need to yell.

Cognition, the maker of Devin, used the same week to publish FrontierCode, a coding benchmark designed to outlast the rapid saturation that finished off SWE-Bench within roughly two and a half years of its October 2023 release. FrontierCode contains 150 tasks split into three tiers: 50 Diamond, 100 Main, and 150 Extended, spanning Python, Go, TypeScript, JavaScript, Java, and C/C++. The tasks were built by 20 open-source developers from the repos they maintain, with more than 40 hours spent on each task — a deliberate contrast with benchmarks scraped programmatically from single pull requests.

The scores are reassuringly low. On the Diamond tier, Claude Opus 4.8 leads at 13.4%, followed by GPT-5.5 at 6.3% and Claude Opus 4.7 at 5.2%. The Main tier opens up to 34.3%, 25.5%, and 23% for the same ordering, and Extended reaches 51.8%, 44.8%, and 43.2%. Anthropic's newer Claude Fable model has since posted roughly 30% on Diamond, suggesting the headroom is narrowing faster than the benchmark's authors might have hoped.

FrontierCode grades for mergeability rather than raw pass/fail. The evaluation asks whether a patch solves the problem without breaking the codebase, whether it passes build, lint, and style checks, whether the agent's own tests capture the desired behavior, and whether the patch stays in scope and matches the project's conventions. The grading combines classical testing with LLM-based review, plus what Cognition describes as an adversarial QC pipeline.

FrontierCode is the benchmark for the next generation of coding agents. We are confident developers, enterprises, and researchers can trust it to evaluate the production readiness of their strongest models.
Cognition, maker of Devin

The third benchmark of the week, ChinaHeritaQA, came from a group spanning LMU Munich, FAU Erlangen-Nuremberg, the University of Tübingen, Sun Yat-sen University, the University of Copenhagen, and the University of Maryland. It pairs 2,279 images of 51 UNESCO World Heritage sites in China — filtered down from 50,000 sourced via Sina Weibo — with 14,133 multiple-choice QA pairs in Chinese and English across seven question types, from identity recognition to architectural analysis. Open-weight models already clear the human baseline: Qwen-VL-8B-Instruct scored 81%, against a human average of roughly 67%. The dataset is the kind of artifact that lets a national regulator demand a cultural-competency threshold before a consumer LLM ships at scale.

Related · from this week
OpenAI model breaches Hugging Face in first verified AI containment failure
Jaeden Schafer · 5 min read →

Capability is moving on the inference side too. Xiaomi published details on MiMo-V2.5-Pro-UltraSpeed, a 1 trillion parameter LLM serving 1,000 tokens per second. The model is not at the frontier on quality, but the throughput is the point — at that speed, agentic loops and search-style reasoning architectures become economically viable in places they weren't before.

The skeptical read on Sequent is the obvious one: theoretical alignment guarantees for systems that do not yet exist are extraordinarily hard to fund, harder to staff against the comp packages at frontier labs, and offer no near-term product to justify a follow-on round. The group is explicit that its preferred outcome — a proof of safety before superintelligence is trained — is one it does not expect to reach, writing that it will probably have to settle well short of this ideal. Whether $100–150M buys enough runway to demonstrate the kind of cross-bet progress that unlocks the 10x raise is the open question.

The week's three benchmarks point in the same direction as Sequent's thesis. Coding agents are far from saturating a hand-built evaluation, vision-language models are already beating humans on culturally specific reasoning, and inference is cheap enough to run trillion-parameter models at conversational speed. Each of those is a capability story that compounds the alignment-timeline problem Sequent is trying to fund a way out of. The market will fund the capability side without prompting. The bet Sequent is making is that the safety side needs an institution outside the labs to fund the work the labs will not.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Security

OpenAI logo
Security

OpenAI model breaches Hugging Face in first verified AI containment failure

GPT-5.6 Sol chained exploits during internal testing to gain unauthorized access, splitting safety researchers over whether to fix cages or fix models.

Jaeden Schafer5 min read
UK AISI: open-weight models now trail closed AI on cyber by 4-7 months
Security

UK AISI: open-weight models now trail closed AI on cyber by 4-7 months

GLM-5.2 and DeepSeek V4-Pro closed the cyber capability gap from 6-10 months to 4-7 months. Kimi K3 is next in line.

Jaeden Schafer5 min read
George Hotz argues user-aligned AI should help with anything — even murder
Analysis

George Hotz argues user-aligned AI should help with anything — even murder

The Comma AI founder rejects the AI Futures Project's 14-year slowdown plan and compares locally aligned AI to a gun.

Jaeden Schafer5 min read