2026-07-30
FLUX 3 Model Overview: Multimodal Flow Models for Image, Video, Audio, and Action PredictionFLUX 3 Model Overview: Multimodal Flow Models for Image, Video, Audio, and Action Prediction
Black Forest Labs 发布多模态流匹配基础模型 FLUX 3,用 Self-Flow 自监督流匹配框架统一图像、视频、音频与动作预测,视频人类偏好率对 Luma Ray 3.2 达 93%,微调后机器人操纵成功率从 42%…FLUX 3 is a diffusion transformer (DiT) trained jointly across images, video, and audio within a unified multimodal architecture.
FLUX 3Black Forest LabsSelf-Flowflow matchingdiffusion transformer
原文Source
2026-07-06
GLM-5.2: Built for Long-Horizon TasksGLM-5.2: Built for Long-Horizon Tasks
Zhipu AI/Z.ai 发布的长程任务旗舰模型 GLM-5.2:首次在稳定 1M-token 上下文上兑现长程 agentic 编码能力,提出 IndexShare 降低稀疏注意力 indexer 开销(1M 长度下单 token F…We're introducing GLM-5.2, our latest flagship model for long-horizon tasks.
GLM-5.2long-horizon agent1M contextDSA sparse attentionIndexShare
原文Source
2026-07-06
GEAR: Guided End-to-End AutoRegression for Image SynthesisGEAR: Guided End-to-End AutoRegression for Image Synthesis
GEAR 让 VQ tokenizer 与自回归生成器联合端到端训练:用 hard/soft 双读出机制绕开 VQ 索引不可微、straight-through estimator 会坍缩的难题,把 ImageNet gFID 收敛速度加…Visual generative models are typically trained in two stages. A tokenizer is first trained for reconstruction and then frozen, after which a generator is train…
autoregressive image generationvector quantizationVQ tokenizerrepresentation alignmentREPA
arXiv:2606.32039
2026-06-19
Flow-GRPO: Training Flow Matching Models via Online RLFlow-GRPO: Training Flow Matching Models via Online RL
Flow-GRPO 首次将在线策略梯度 RL(GRPO)引入 flow matching 文本到图像模型,通过 ODE-to-SDE 转换注入随机性,并采用 Denoising Reduction 加速训练采样,使 SD3.5-M 的 G…We propose Flow-GRPO, the first method to integrate online policy gradient reinforcement learning (RL) into flow matching models.
flow matchingGRPO在线强化学习ODE-to-SDEreward hacking
arXiv:2505.05470
2026-06-14
MeshFlow: Efficient Artistic Mesh Generation via MeshVAE and Flow-based Diffusion TransformerMeshFlow: Efficient Artistic Mesh Generation via MeshVAE and Flow-based Diffusion Transformer
MeshFlow 提出 MeshVAE(连续隐空间压缩网格几何与拓扑)+ Rectified Flow DiT(并行去噪所有 latent token),约 1.2 秒生成 artist-like 三角网格,比最快 AR 方法快 18 倍…We present MeshFlow, a new method for generating artist-like 3D meshes. Current mesh generators often adopt Auto-Regressive (AR) next-token prediction, a natur…
mesh generationMeshVAEflow matchingdiffusion transformer三维网格生成
arXiv:2606.04621
2026-06-10
Yume: An Interactive World Generation ModelYume: An Interactive World Generation Model
Yume 提出量化相机运动、Masked Video Diffusion Transformer、无训练 Anti-Artifact Mechanism 与 TTS-SDE 采样器四大组件,以单张图像为输入、键盘为控制接口,自回归生成理论…Yume aims to use images, text, or videos to create an interactive, realistic, and dynamic world, which allows exploration and control using peripheral devices…
interactive world generationvideo diffusion量化相机运动 QCMMasked Video Diffusion Transformerautoregressive video
arXiv:2507.17744
2026-06-10
WorldMem: Long-term Consistent World Simulation with MemoryWorldMem: Long-term Consistent World Simulation with Memory
WorldMem 通过引入记忆库与状态感知 cross-attention,让视频扩散模型在生成超长序列时仍能忠实重建先前观测的场景,在 Minecraft 超窗口基准上 PSNR 较基线提升 6.66,LPIPS 降低至三分之一。World simulation has gained increasing popularity due to its ability to model virtual environments and predict the consequences of actions.
world simulationmemory banklong-term consistencyvideo diffusion状态感知注意力
arXiv:2504.12369
2026-06-10
Towards Uncertainty Quantification in Generative Model LearningTowards Uncertainty Quantification in Generative Model Learning
本文首次形式化生成模型评估中的不确定性量化问题,提出基于集成 Precision-Recall 曲线的方法,通过多次随机初始化训练的模型集合来捕获学习分布近似目标分布时的置信区间,并在合成 DDPM 实验中验证了该方法可有效揭示模型复杂度…While generative models have become increasingly prevalent across various domains, fundamental concerns regarding their reliability persist.
不确定性量化Generative ModelsPrecision-Recall CurvesEpistemic UncertaintyEnsemble Methods
arXiv:2511.10710
2026-06-10
The Matrix: Infinite-Horizon World Generation with Real-Time Moving ControlThe Matrix: Infinite-Horizon World Generation with Real-Time Moving Control
首个可在实时交互控制下生成无限长高保真720p视频流的世界模拟器,结合AAA级游戏训练数据(Forza Horizon 5、Cyberpunk 2077)、Swin-DPM滑动窗口无限生成与Stream Consistency Model…We present The Matrix, the first foundational realistic world simulator capable of generating continuous 720p high-fidelity real-scene video streams with real-…
world modelvideo generationreal-time controlinfinite-horizon世界模型
arXiv:2412.03568
2026-06-10
Quantifying Epistemic Uncertainty in Diffusion ModelsQuantifying Epistemic Uncertainty in Diffusion Models
FLARE 方法通过 Fisher 信息将认知不确定性从扩散模型的随机采样噪声中显式分离,在三个合成时间序列基准上以最高 93.08% Gap-Closure 大幅优于 BayesDiff 和 LLLA 等现有方法。To ensure high quality outputs, it is important to quantify the epistemic uncertainty of diffusion models.
扩散模型Epistemic UncertaintyLaplace ApproximationFisher InformationFLARE
arXiv:2602.09170
2026-06-10
PlayerOne: Egocentric World SimulatorPlayerOne: Egocentric World Simulator
PlayerOne 是首个以真实人体运动(SMPL)为条件的第一人称世界模拟器,基于 Diffusion Transformer,通过 Part-Disentangled Motion Injection 和 Scene-Frame Re…We introduce PlayerOne, the first egocentric realistic world simulator, facilitating immersive and unrestricted exploration within vividly dynamic environments.
egocentric world simulator自我中心视角生成diffusion transformermotion injectionSMPL 人体姿态
arXiv:2506.09995
2026-06-10
Playable Video GenerationPlayable Video Generation
CADDY 在完全无标注视频上自监督学习离散动作空间,让用户像玩游戏一样逐帧控制视频生成,CVPR 2021 Oral,在 BAIR、Atari Breakout 和 Tennis 三个数据集上全面超越基线。This paper introduces the unsupervised learning problem of playable video generation (PVG). In PVG, we aim at allowing a user to control the generated video by…
playable video generationunsupervised action learning可交互视频生成discrete action spaceCADDY
arXiv:2101.12195
2026-06-10
PhyCo:面向生成式运动的可控物理先验学习PhyCo: Learning Controllable Physical Priors for Generative Motion
PhyCo 通过构建大规模物理仿真数据集、ControlNet 像素对齐属性条件调节与 VLM 奖励优化三者结合,实现了对摩擦、弹性、形变及外力等物理属性的连续可控视频生成,在 Physics-IQ 基准上达到新 SOTA。Modern video diffusion models excel at appearance synthesis but still struggle with physical consistency: objects drift, collisions lack realistic rebound, and…
物理先验视频生成ControlNetVLM奖励优化物理仿真数据集
arXiv:2604.28169
2026-06-10
Nano World Models: A Minimalist Implementation of Future Video PredictionNano World Models: A Minimalist Implementation of Future Video Prediction
NanoWM 是以 diffusion forcing 为核心的极简视频预测世界模型框架,通过统一接口系统研究预测目标、模型规模、动作注入与潜在空间对视频预测质量和长程自回归行为的影响,并完整开源代码、权重与数据。World models have become a central paradigm for learning predictive simulators that support generation, planning, and decision-making.
world modeldiffusion forcing视频预测action conditioninglong-horizon rollout
arXiv:2605.23993
2026-06-10
Matrix-Game: Interactive World Foundation ModelMatrix-Game: Interactive World Foundation Model
Matrix-Game 是一个 170 亿参数的交互式游戏世界基础模型,通过两阶段训练(无标注视频预训练 + 精标动作微调)在 Minecraft 上实现精确的键盘/鼠标帧级控制与高质量视频生成,并提出 GameWorld Score 统…We introduce Matrix-Game, an interactive world foundation model for controllable game world generation.
world modelinteractive generationdiffusion transformeraction controllabilityautoregressive generation
arXiv:2506.18701
2026-06-10
Improved Mean Flows:加速生成模型的挑战与改进Improved Mean Flows: On the Challenges of Fastforward Generative Models
本文提出 iMF,通过将 MeanFlow 训练目标重构为瞬时速度损失、引入灵活的 CFG 条件化和轻量级 in-context conditioning,在 ImageNet 256×256 上实现单步(1-NFE)FID 1.72,无…MeanFlow (MF) has recently been established as a framework for one-step generative modeling. However, its ``fastforward'' nature introduces key challenges in b…
MeanFlowflow matching单步生成classifier-free guidancevelocity loss
arXiv:2512.02012
2026-06-10
Genie: Generative Interactive EnvironmentsGenie: Generative Interactive Environments
Genie 是首个从无标注互联网视频无监督训练的 110 亿参数生成式交互环境基础模型,能够从单张图片或文字提示生成可逐帧交互的虚拟世界,并自动学习离散潜在动作空间以支持智能体训练。We introduce Genie, the first generative interactive environment trained in an unsupervised manner from unlabelled Internet videos.
world modelgenerative interactive environmentlatent action modelspatiotemporal transformerMaskGIT
arXiv:2402.15391
2026-06-10
Generative Uncertainty in Diffusion ModelsGenerative Uncertainty in Diffusion Models
提出基于 Laplace 近似的贝叶斯框架,通过语义似然度量化扩散模型每个生成样本的生成不确定性,自动过滤低质量图像,在 ImageNet 上将 UViT 的 FID 从 9.45 提升至 7.89。Diffusion models have recently driven significant breakthroughs in generative modeling. While state-of-the-art models produce high-quality samples on average,…
Diffusion ModelsGenerative UncertaintyBayesian InferenceLaplace Approximation不确定性估计
arXiv:2502.20946
2026-06-10
GameGen-X: Interactive Open-world Game Video GenerationGameGen-X: Interactive Open-world Game Video Generation
GameGen-X 是首个专为开放世界游戏视频生成与交互控制设计的 Diffusion Transformer 模型,通过两阶段训练(基础模型预训练 + InstructNet 指令微调)和百万级 OGameData 数据集,实现高质量游…We introduce GameGen-X, the first diffusion transformer model specifically designed for both generating and interactively controlling open-world game videos.
game video generationdiffusion transformeropen-world gameinteractive controlInstructNet
arXiv:2411.00769
2026-06-10
GameFactory: Creating New Games with Generative Interactive VideosGameFactory: Creating New Games with Generative Interactive Videos
GameFactory 利用预训练视频扩散模型的开放域生成先验,通过 domain adapter 与四阶段多阶段训练策略将 game style 学习与 action control 解耦,实现跨场景可泛化的键鼠动作控制游戏视频生成,并…Generative videos have the potential to revolutionize game development by autonomously creating new content.
游戏视频生成video diffusionaction controlscene generalizationstyle-action decoupling
arXiv:2501.08325
2026-06-10
Diffusion Model Guided Sampling with Pixel-Wise Aleatoric Uncertainty EstimationDiffusion Model Guided Sampling with Pixel-Wise Aleatoric Uncertainty Estimation
一种无需训练的扩散模型逐像素 aleatoric uncertainty 估计方法,通过对去噪分数施加扰动并计算方差得到不确定性图,并将其用于引导采样,在 ImageNet 和 CIFAR-10 上以仅 20 NFEs(比 BayesDi…Despite the remarkable progress in generative modelling, current diffusion models lack a quantitative approach to assess image quality.
diffusion modelaleatoric uncertaintypixel-wise uncertaintyguided samplingFID
arXiv:2412.00205
2026-06-10
扩散作为自蒸馏:单模型端到端潜变量扩散Diffusion As Self-Distillation: End-to-End Latent Diffusion In One Model
本文提出 Diffusion as Self-Distillation(DSD)框架,通过 Stop-Gradient 解耦、损失变换和 EMA 目标编码器三项设计,将 VAE 编解码器与扩散网络统一为单一可训练模型,解决联合训练中的潜空…Standard Latent Diffusion Models rely on a complex, three-part architecture consisting of a separate encoder, decoder, and diffusion network, which are trained…
latent diffusion modelend-to-end trainingself-distillation潜空间坍塌VAE联合训练
arXiv:2511.14716
2026-06-10
Cosmos World Foundation Model Platform for Physical AICosmos World Foundation Model Platform for Physical AI
NVIDIA Cosmos 是面向 Physical AI 的开源世界基础模型平台,涵盖视频数据处理飞轮、高效视频分词器(连续/离散,速度快 2-12×)、7B/14B 扩散式与 4B-13B 自回归式预训练 WFM,以及机器人操作、自动…Physical AI needs to be trained digitally first. It needs a digital twin of itself, the policy model, and a digital twin of the world, the world model.
world foundation modelPhysical AIvideo generationvideo tokenizerdiffusion model
arXiv:2501.03575
2026-06-10
BayesDiff: Estimating Pixel-wise Uncertainty in Diffusion via Bayesian InferenceBayesDiff: Estimating Pixel-wise Uncertainty in Diffusion via Bayesian Inference
BayesDiff 通过 Last-Layer Laplace Approximation 与 Uncertainty Iteration Principle,在扩散模型反向生成链中估计逐像素不确定性,实现低质量图像过滤、多样性增强与 a…Diffusion models have impressive image generation capability, but low-quality generations still exist, and their identification remains challenging due to the…
Diffusion ModelsUncertainty QuantificationBayesian InferenceLaplace ApproximationImage Generation Quality
arXiv:2310.11142
2026-06-10
Adversarial Flow Models — 论文海报Adversarial Flow Models
本文提出 Adversarial Flow Models(AFM),通过在对抗目标上叠加最优传输正则化损失并引入梯度归一化技术,将 GAN 与 Flow Matching 融合,在 ImageNet 256px 单步图像生成上以 112…We present adversarial flow models, a class of generative models that belongs to both the adversarial and flow families.
adversarial trainingflow matchingoptimal transport单步图像生成GAN
arXiv:2511.22475