10 Oct 2026
OllyGarden raises $4m seed to clean up OpenTelemetry data
Founded in 2025 by OpenTelemetry contributors, the company launched Minimum Viable Instrumentation with the round. Next Frontier Capital,…
Tencent Youtu Lab's 5B parsing model, released 9 October 2026: turns documents, charts, audio and video into one structured JSON output. Open weights under a custom licence that excludes use within the European Union.
Not yet independently verified. The release date of 9 October 2026 rests only on the creation of the Hugging Face repository; the model card dates its report and weights only as October 2026 and lists the report as coming soon. The benchmark figures are Tencent's own, from the model card, so no score is stored. The licence is a custom one that states the model is not intended for use within the European Union. No independent report was found. We will update this when it can be confirmed, and remove this note.
Youtu-Parsing-Omni is a 5 billion parameter model from Tencent that turns documents, images, charts, geometry figures, audio and video into structured JSON. Its Hugging Face repository was created on 9 October 2026. Its licence states that it is not intended for use within the European Union.
A single encoder feeds a Youtu-LLM decoder, and the output is one JSON block covering perception (layout, text, tables, formulas, bounding boxes, timestamps, speech recognition and character recognition) and cognition (captions, narratives and reports). The card lists task modes for documents, natural images, charts, flowcharts, geometric figures, audio, natural video and text-rich video. Document output includes LaTeX formulas, tables, Markdown and Mermaid diagrams. Bounding boxes use a 0 to 1000 grid and timestamps use hours, minutes and seconds.
On OmniDocBench v1.6 the card reports an overall score of 96.96, the top of its table. These are Tencent's own results, and the card shows other systems ahead on some tests, for example average similarity on ChemOCR.
The model suits developers and researchers who run their own document and media parsing pipelines and have CUDA GPUs. Serving uses vLLM with an OpenAI-compatible endpoint; audio and video input needs ffmpeg. Loading through Transformers runs code shipped with the checkpoint, so the card advises pinning a reviewed revision.
| Benchmark | Official | Community avg |
|---|---|---|
| No benchmark scores yet. Be the first to add one. | ||
“Official” values are editor-approved and feed the ranking. “Community avg” is the mean of member submissions (shown for transparency; it never affects the ranking until an editor approves a value).
Sign in to add a benchmark score for this model.
An earned signal from verification, reviews, awards, transparency and engagement - the vendor can't buy it.
Updated 10/10/2026
Used it? Your experience helps other buyers decide.
Write a reviewNo questions yet. Be the first to ask about Youtu-Parsing-Omni.
Everything here links back to the same verified catalogue. Pick your next stop.