Skip to content
TrustList
Funding news

Clockwork.io raises 31 million dollars from Premji Invest, Wing and Seligman for AI fault-tolerance software

Editorial

By TrustList Editorial

Clockwork.io says it has raised 31 million dollars, co-led by Seligman Ventures, Wing Ventures and Premji Invest, to extend its software that keeps AI training and inference jobs running through GPU and network failures.

About Clockwork.io raises 31 million dollars from Premji Invest, Wing and Seligman for AI fault-tolerance software

Clockwork.io raises 31 million dollars from Premji Invest, Wing and Seligman for AI fault-tolerance software

5 October 2026 — Clockwork.io, a Palo Alto software company, has announced a 31 million dollar funding round to expand software that keeps large AI training and inference workloads running when hardware fails. The company's release was issued on 5 October.

Not yet independently verified. The release was read only through a wire edition and a search summary; the company's own site did not carry the announcement on its blog at the time of reading. SiliconANGLE calls the round Series C while FinSMEs gives no stage, so the stage is left out of the headline and should be confirmed. The 73 million dollar total is from press. Customer deployments are the company's own claims. We will update this when it can be confirmed, and remove this note.

The round

The company states the round as 31 million dollars. It was co-led by Seligman Ventures, Wing Ventures and Premji Invest, with existing investors New Enterprise Associates and e& Capital returning. Press coverage describes it as a Series C and puts total funding at 73 million dollars, but the stage is not confirmed from the company's own wording and TrustList therefore does not use it in the title. The company says the money will support broader enterprise adoption and delivery through cloud partners, and the release was published alongside a new feature called TorchSnap.

What the company does

Clockwork sells software that sits between AI workloads and GPU clusters. Its products include TorchPass and LinkPass, which handle failover, migration of work between GPUs and the capture of snapshots and checkpoints so a job can restart with little lost progress. TorchSnap is described as capturing multi-node snapshots of distributed inference jobs without code changes. The release names LinkedIn, Together AI and WhiteFiber as users, with LinkedIn said to have avoided thousands of GPU-hours of downtime each month. Those are the company's claims.

Why it matters for buyers

GPU time is one of the largest costs in running AI at scale, and failures that force a job to restart waste it. Organisations renting or operating clusters, and cloud providers that sell them, are the target buyers. The round shows continued investor interest in the reliability layer rather than in models themselves. Buyers should note that the product is aimed at large distributed workloads and may be of limited use for small deployments.

What to check

  • Whether the software works with your training and inference frameworks and your cluster scheduler, and what code changes if any are needed.
  • Overhead: the cost in performance of continuous checkpointing and snapshots.
  • Whether it is sold directly or through a cloud provider, and the pricing basis.
  • How snapshots are stored, encrypted and deleted, since they may contain model weights or customer data.
  • Independent evidence for the downtime-saving claims, which come from the company's own release.

A pilot on a representative workload is the most reliable test before committing.

Company profile on TrustList: Clockwork.io

Sources

Categories & features