arxiv:2602.02493
ZehongMa
zehongma
AI & ML interests
MLLMs, Image/Video Generation, Multi-modal Representation Learning
Recent Activity
authored a paper about 21 hours ago
PixelGen: Pixel Diffusion Beats Latent Diffusion with Perceptual Loss upvoted a paper about 1 month ago
UniDDT: Unifying Multimodal Understanding and Generation with Decoupled Diffusion Transformer upvoted a paper about 2 months ago
Representation Forcing for Bottleneck-Free Unified Multimodal ModelsOrganizations
None yet