Blacksmith has raised a $45 million Series B led by Peak XV Partners at a $550 million valuation, a nearly 10x jump from the $60 million mark set when it closed a $10 million Series A less than a year ago. GV and Y Combinator followed on, bringing total funding to $58.5 million. The pitch is straightforward: AI is generating far more code than human reviewers and existing test pipelines can validate, and Blacksmith sells the infrastructure to catch up.
Founded in 2024, Blacksmith runs the continuous integration workloads that companies use to build, test, and verify software before shipping. Customer count has climbed from more than 700 less than a year ago to more than 5,000 today, with Mercury, Supabase, Clerk, Ashby, and Expensify among the names. Co-founder and CEO Aditya Jayaprakash said the company hit a $10 million annualized run rate with just 10 employees and has since scaled to about 30 staff and 'tens of millions' in revenue. He declined to give a precise updated ARR but said the largest customers now spend more than $1 million a year on the platform.
The tailwind is the surge in AI-written code from tools like Cursor, OpenAI's Codex, and Anthropic's Claude Code. Faster generation has not come with faster review; if anything, the volume has widened the gap between code produced and code verified.
Key facts
- 01Blacksmith raised a $45M Series B led by Peak XV Partners at a $550M valuation, up from $60M less than a year ago.
- 02Customer count grew from 700 to more than 5,000 in under a year, with Mercury, Supabase, Clerk, Ashby, and Expensify on the roster.
- 03The company hit $10M ARR with just 10 employees and now runs 30 staff at 'tens of millions' in revenue.
- 04Largest customers now spend more than $1M annually on the platform.
- 05Total funding to date reaches $58.5M, with GV and Y Combinator following on.
That framing — validation as the new bottleneck — is the entire investment thesis behind the round. Every additional AI coding seat sold by a rival lab generates more downstream demand for the CI runs, test executions, and automated fix cycles Blacksmith charges for. The startup has extended beyond raw CI capacity with Codesmith, an AI coding agent that automatically fixes failed checks, moving the product from passive test runner to active repair layer.
The competitive picture is crowded. GitHub Actions remains the default CI runner for most teams, Cursor Automations pushes similar workflows from the IDE side, and Codex and Claude Code are folding in their own validation loops. Amazon Web Services, Microsoft Azure, and Google Cloud all offer AI-flavored code testing services, and a long tail of startups is chasing the same category. Jayaprakash said Blacksmith is competing on the speed of test runs and pricing, two axes where hyperscaler services tend to be slowest to move.
Speed matters more than it used to. A test suite that took ten minutes when a human wrote one pull request per day becomes a serious constraint when an AI agent opens twenty pull requests per hour. Teams working with autonomous coding agents need CI that returns results in seconds, not minutes, or the agent stalls waiting for feedback. That is the wedge Blacksmith is arguing separates it from GitHub Actions and the hyperscaler offerings.
The valuation math is aggressive. At $550 million against 'tens of millions' in revenue, the multiple sits in the range that only holds up if AI-coding-driven CI demand keeps compounding. Peak XV, GV, and Y Combinator are betting that the shift from human-authored to AI-authored code is a durable secular trend, not a 2025 novelty.
Risk sits on two axes. First, the foundation model providers keep absorbing adjacent categories — Codex and Claude Code already run tests as part of their agent loops, and it is not obvious that a standalone CI provider stays essential if those loops get good enough. Second, GitHub Actions has the distribution advantage of being wired into every repo on the largest code host, which is a hard moat to unseat on pricing alone.
Jayaprakash said Blacksmith plans to expand into a broader suite of coding tools to help developers write, validate, and merge software faster — a signal that the company sees itself moving up-stack toward the same territory as Cursor and Codex, not just serving them.
The interesting read on this round is what it says about where AI coding revenue actually accumulates. Frontier labs get the headlines and the biggest checks, but the picks-and-shovels layer — the CI capacity, the test infrastructure, the automated repair agents — is where a lot of the operating spend from AI-native engineering teams is landing. If Blacksmith's ARR trajectory holds, it is early evidence that validation infrastructure is becoming a meaningful line item in engineering budgets, not a rounding error on the model bill. That is a durable place to sit, provided the hyperscalers and the model labs do not decide to give it away.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




