
NVIDIA PixelUMM Shipped Without an Announcement — The Weights Just Landed
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 219 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAINEWGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 117 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1064 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 41 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 104 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 213 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
The easiest way to find out that NVIDIA built a new unified multimodal model is to notice that nobody told you. The nvidia/PixelUMM repository appeared on Hugging Face at 21:40 UTC on 1 October 2026, carrying a 15.2-billion-parameter checkpoint whose entire learned stack sits on top of a Qwen3-8B backbone — and there was no blog post, no press release, no keynote slide and no launch thread to go with it. Searches for the name turn up nothing on NVIDIA's own blog. What exists instead is a GitHub repository, a project page whose own URL still says "preview", an arXiv preprint numbered 2609.38597, and a checkpoint split across 128 files with a hidden index that the loader refuses to run without. PixelUMM is real and downloadable today. Whether NVIDIA considers it released is a question the company has not answered.
That gap — between an artifact that exists and a vendor that has said nothing — is the whole story here, and it is worth being precise about which side of it each fact sits on. Everything below comes from the repository, the model card and the paper the authors themselves published. Nothing comes from an announcement, because there wasn't one.
What appeared, and when
The trail runs about four weeks, and each piece landed quietly.
• 4 September 2026 — nv-tlabs/PixelUMM is created on GitHub under an Apache-2.0 repository license, described simply as "Encoder-Free Unified Image and Video Understanding and Generation". It sits with a single initial commit until the end of September.
• 28–29 September 2026 — the repository's initial commit lands, and arXiv assigns the preprint arXiv:2609.38597, dated 29 September, listing authors from NVIDIA and the University of Waterloo: Cong Wei, Xuanchi Ren, Bryan Chu, Weiming Ren, Huan Ling, Jiahui Huang, Laura Leal-Taixé, Sanja Fidler, Wenhu Chen, Zian Wang and Jay Zhangjie Wu.
• 1 October 2026, 21:32 UTC — a final commit to the repository titled "docs: add PixelUMM paper citation".
• 1 October 2026, 21:40 UTC — the Hugging Face model repository is created, eight minutes later, and populates with the checkpoint shards.
The project page is hosted at a path that literally reads pixelumm-project-page-preview, and the paper carries no venue line — no CVPR, no NeurIPS, no "accepted to". Taken together the sequence reads like a research team pushing a paper artifact live and letting the HF upload be the announcement. That is an observation about the evidence, not a claim about intent: NVIDIA may well have a launch planned for later, and nothing here rules that out.

What PixelUMM actually is
The design thesis is stated in the first line of the model card: no VAE, no vision encoder. Where a conventional unified model carries two visual interfaces — a vision transformer producing semantic features for understanding, and a variational autoencoder producing reconstruction latents for generation — PixelUMM carries none. Images are cut into 16×16 pixel patches; videos are cut into 4-frame spatiotemporal tubelets; both reach the backbone through nothing more than single-layer linear projections. Raw pixels go in; raw pixels come out.
• Architecture — decoder-only Transformer with raw-pixel patch embeddings and an iterative pixel-generation head, described in the paper as a Mixture-of-Transformers that pairs shared attention with task-specific parameters.
• Backbone — Qwen3-8B, pinned to revision b968826d9c46dd6066d109eabc6255188de91218. Only its config and tokenizer files are needed; the checkpoint carries its own learned language weights.
• Parameters — 15,199,672,064 (roughly 15.2B), rendered as "8B MoT" in the paper's own benchmark tables, where the size notation counts the backbone and generation expert separately.
• Objectives — autoregressive text prediction and pixel-space flow matching trained jointly, which is what lets one set of weights both answer a question about an image and draw a new one.
• Tasks — text-to-image, text-to-video at 96 frames / 24 fps / 4 seconds, image-conditioned text, and video-conditioned text.
The paper's empirical section is unusual in a way that is worth flagging. Eight of its sections are studies of design choices rather than leaderboard placements — image patch size, video patch size, patch artifacts, pixel-space versus VAE-space training dynamics, model size, compute scaling, multimodal context conditioning, and video understanding interfaces. That is a paper written by people trying to answer whether the approach works, not one trying to win a table.
The numbers are the authors' own
Every figure below is reported by the PixelUMM authors in their own preprint. There is no independent reproduction, no arena Elo, and no third-party evaluation, because the model is six weeks old at most and predates any external harness. Treat these as claims with a download link, not as verified performance.
• Image understanding — MMMU 41.67, MMStar 53.99, AI2D 80.12, DocVQA 90.42, ChartQA 82.96, OCRBench 78.00, BLINK 53.46, MMMU-Pro 27.63, over the official LMMS-Eval protocol (64,750 generations across 21 tasks).
• Video understanding — MVBench 70.53, Video-MME 57.33 without subtitles, LongVideoBench 59.61, LVBench 40.41.
• Image generation — GenEval overall 0.83 with an LLM prompt rewriter, 0.77 without one; DPG-Bench overall 85.74.
• Video generation — VBench Part 1 quality score 84.10, semantic score 79.80; VBench Part 2 total 83.24.
The honest reading is that these are competitive-within-class numbers, not category-leading ones, and the paper says so itself: it notes that because training data differ across models, the results "cannot establish which architecture is superior". Against Qwen3-VL-8B the image-understanding gap is large on MMMU (41.67 against 69.60) and on MMMU-Pro where the authors report no competing figure. Against Qwen-Image 20B the GenEval gap is 0.83 to 0.87. What PixelUMM is not, on its own numbers, is a state-of-the-art model. What it is, on its own numbers, is a 15B model with a genuinely unusual architecture that lands in the same band as the specialist systems around it.
The sharpest edge is the license, not the benchmark
The repository is Apache-2.0 and the model card advertises that plainly. The checkpoint is a different artifact with different terms, and this is the detail most likely to catch a team out. The weights ship under the NVIDIA One-Way Noncommercial License, whose use is limited to non-commercial research or evaluation — a materially tighter grant than the code sitting next to it. One source file in the repository, modeling/pixelumm/modeling_utils.py, additionally retains a CC BY-NC 4.0 notice derived from DiT.
For a research group, an evaluation team or anyone publishing a paper, this is a perfectly usable license. For a product team prototyping an internal feature, it is the first thing to put in front of legal, and the answer may be no. A model you cannot ship is a different kind of asset from a model you can, and no amount of benchmark parity changes that.

What it takes to run it
This is not a weekend download. The environment documentation asks for Linux x86-64, Python 3.12, a CUDA 13.0 development toolkit, an NVIDIA GPU, and FlashAttention built from source against that toolkit — a runtime-only CUDA image will not do it. The checkpoint itself is not a model.safetensors file: it is 128 .distcp shards totalling roughly 30 GB plus a hidden .metadata index, and the loader requires every referenced shard and that index to be present.
Two further details shape what you can actually do with it. First, text-to-video runs Cosmos guardrails by default, and those need access to the gated nvidia/Cosmos-1.0-Guardrail repository plus a second Python environment with a different Transformers major version — a login alone does not unlock the weights. Second, the four-step toy training example, the only training recipe published, is documented as needing seven GPUs with at least 48 GiB each and about 61 GB of free disk for the output. Fine-tuning on a workstation is not the intended path; inference on a single modern card is.
The repository ships four checkpoints. S8-F22-R05 is the default and the one used for the paper's evaluation, covering all four tasks. S8-F18-R01 took 10,000 additional fine-tuning steps at 480p and 720p and generally gives slightly better text-to-video results, but it cannot do video understanding. S8-F19-R03 and S8-F21-R02 are intermediate stages. Choosing between them is a real decision, not a detail.
What is not confirmed
A short list, and it matters more than the long one above.
• No NVIDIA announcement. No press release and no post on the company's own blog surfaced for this model at the time of writing. The quiet release may be deliberate — research artifacts often ship this way — or a launch may simply not have happened yet.
• No independent evaluation. Every number in this article is the authors' own. Nothing has been rerun outside NVIDIA and Waterloo.
• No hosted route anywhere. You cannot call PixelUMM through an API today, and that includes OrcaRouter — we do not route it, and cannot, because a noncommercially licensed checkpoint is not something a commercial serving platform can offer. Anyone who tells you otherwise is describing a self-hosted setup.
• No stated position on future licensing. The noncommercial terms are the terms as shipped. Whether they relax is unknown and, given NVIDIA's history with research checkpoints, not something to plan around.

What to watch for next
Three events would turn PixelUMM from a research artifact into something a broader audience can use. The first is a licensing change on the checkpoint — that single file is the entire barrier between "interesting" and "usable in a product". The second is an official NVIDIA launch, which would come with a framing the repository cannot supply: what the model is for, and whether it is a product direction or a paper. The third is the first independent reproduction, most likely a GenEval or MVBench rerun by someone with a spare GPU cluster, which is the moment the authors' numbers stop being the only numbers.
Until then the correct posture is the one the evidence supports. PixelUMM exists, the code and paper are public and readable, the weights download, and none of the performance claims have been checked by anyone outside the team that made them. That is not a criticism of the work — it is what a six-week-old research drop looks like. It is also exactly the situation where a routing layer earns its keep later, if the license ever loosens: one key across the models you already trust, list prices passed through at 0% markup, and failover that lets you point a slice of traffic at something unproven without betting a production path on it. For now, though, the honest summary is simpler. NVIDIA built something unusual, published it thoroughly and told nobody. The repository is the announcement.
