Skip to content
TrustList
YP
AI Model

Youtu-Parsing-Omni

New· 4

Tencent Youtu Lab's 5B parsing model, released 9 October 2026: turns documents, charts, audio and video into one structured JSON output. Open weights under a custom licence that excludes use within the European Union.

About Youtu-Parsing-Omni

Not yet independently verified. The release date of 9 October 2026 rests only on the creation of the Hugging Face repository; the model card dates its report and weights only as October 2026 and lists the report as coming soon. The benchmark figures are Tencent's own, from the model card, so no score is stored. The licence is a custom one that states the model is not intended for use within the European Union. No independent report was found. We will update this when it can be confirmed, and remove this note.

Youtu-Parsing-Omni is a 5 billion parameter model from Tencent that turns documents, images, charts, geometry figures, audio and video into structured JSON. Its Hugging Face repository was created on 9 October 2026. Its licence states that it is not intended for use within the European Union.

What Youtu-Parsing-Omni does

A single encoder feeds a Youtu-LLM decoder, and the output is one JSON block covering perception (layout, text, tables, formulas, bounding boxes, timestamps, speech recognition and character recognition) and cognition (captions, narratives and reports). The card lists task modes for documents, natural images, charts, flowcharts, geometric figures, audio, natural video and text-rich video. Document output includes LaTeX formulas, tables, Markdown and Mermaid diagrams. Bounding boxes use a 0 to 1000 grid and timestamps use hours, minutes and seconds.

On OmniDocBench v1.6 the card reports an overall score of 96.96, the top of its table. These are Tencent's own results, and the card shows other systems ahead on some tests, for example average similarity on ChemOCR.

Who it is for

The model suits developers and researchers who run their own document and media parsing pipelines and have CUDA GPUs. Serving uses vLLM with an OpenAI-compatible endpoint; audio and video input needs ffmpeg. Loading through Transformers runs code shipped with the checkpoint, so the card advises pinning a reviewed revision.

Key facts

  • Maker: Tencent
  • Repository created: 9 October 2026
  • Licence: the Youtu-Parsing licence. Its first clause says the model is not intended for use within the European Union and that this clause prevails in any conflict
  • Size: 5B parameters, BF16 weights
  • Pricing: open weights; no price on the model card
  • Platforms: Python 3.10 or later with CUDA GPUs
  • Weights: Hugging Face, tencent/Youtu-Parsing-Omni

Benchmarks & AI stats

BenchmarkOfficialCommunity avg
No benchmark scores yet. Be the first to add one.

“Official” values are editor-approved and feed the ranking. “Community avg” is the mean of member submissions (shown for transparency; it never affects the ranking until an editor approves a value).

Sign in to add a benchmark score for this model.

Trust Score

4/ 100
Trust Score: New

An earned signal from verification, reviews, awards, transparency and engagement - the vendor can't buy it.

Verification
0/100 · 20%
Reviews
0/100 · 30%
Awards
0/100 · 15%
Transparency
18/100 · 20%
Engagement
0/100 · 15%
Joining soon
Recommendations
Coming soon
Complaints
Coming soon

Updated 10/10/2026

Request a demo or quote from Youtu-Parsing-Omni

Protected by reCAPTCHA - Google Privacy Policy and Terms apply.

By sending, you agree we may share your request and contact details with the provider once you confirm your email.

Reviews

Write the first review of Youtu-Parsing-Omni

Used it? Your experience helps other buyers decide.

Write a review

Questions & answers

No questions yet. Be the first to ask about Youtu-Parsing-Omni.