Skip to content
TrustList
D1
AI Model

Darwin-180B-RSI

VIDRAFT’s open 180B reasoning model: Alibaba’s Qwen3.8-Flash-Next retrained on its own checked answers, reading text and images, with a built-in confidence score that is returned before it answers.

About Darwin-180B-RSI

Darwin-180B-RSI is an open-weight reasoning model from VIDRAFT, a Korean AI company, published on Hugging Face on 28 September 2026 and reported the same day by ZDNet Korea and AI Times. It is a fine-tune of Alibaba’s Qwen3.8-Flash-Next, a 180-billion-parameter mixture of experts with 10 of 512 experts active per token, and it keeps that model’s vision encoder and its context of about 262,000 tokens. VIDRAFT calls its method recursive self-improvement: the model solves practice problems, only the answers that check out against an answer key are kept, and it is retrained on that reasoning. VIDRAFT’s own comparison with the parent model shows the change is small in accuracy (88.12 against 88.04 per cent on MMLU-Pro) but cuts average reasoning length by about 11 per cent, which makes it cheaper to run. It also ships a “zero-token confidence” readout that estimates, before answering, how likely the answer is to be right, so an application can escalate or decline instead. VIDRAFT reports 100 on AIME 2026 and HMMT February 2026, 94.44 on GPQA Diamond, 88.12 on MMLU-Pro and 79.48 on MMMU-Pro, and calls these first place on five Hugging Face leaderboards. Those places are self-reported entries that each publisher submits, most of VIDRAFT’s figures are majority votes over up to 16 attempts with a 131,072-token thinking budget, and other models’ entries use different settings, so they are not stored as scores and are not a like-for-like ranking. No independent evaluation had been published when this was read on 28 September 2026. The weights carry the Qwen Community License 1.0, which a business should read before commercial use; the model card describes self-hosting with vLLM and names no hosted API. The sources disagree on its size: the model card gives 180 billion parameters, while ZDNet Korea describes it as a 397-billion-class model; this listing follows the card.

Benchmarks & AI stats

BenchmarkOfficialCommunity avg
No benchmark scores yet. Be the first to add one.

“Official” values are editor-approved and feed the ranking. “Community avg” is the mean of member submissions (shown for transparency; it never affects the ranking until an editor approves a value).

Sign in to add a benchmark score for this model.

Request a demo or quote from Darwin-180B-RSI

Protected by reCAPTCHA — Google Privacy Policy and Terms apply.

By sending, you agree we may share your request and contact details with the provider once you confirm your email.

Reviews

Write the first review of Darwin-180B-RSI

Used it? Your experience helps other buyers decide.

Write a review

Questions & answers

No questions yet. Be the first to ask about Darwin-180B-RSI.