๐Ÿ‘€ About Me

Hi,

๐ŸŒฑ Iโ€™m Xiaoxiao Ma, a second-year PhD student at USTC, in USTC-BIVLab supervised by Prof. Feng Zhao. I am currently a research intern at JD Explore Academy.

๐Ÿ“– My research interests include:

  • Multimodal learning and foundation models
  • Visual generation, including image, video, and audio-visual generation

๐Ÿ“ซ Looking forward to any collaborations or internship positions, feel free to contact me via email



๐Ÿ”ฅ News

  • 2026.07: ย  Fast-ARDiff was accepted to ECCV 2026 as a Spotlight!
  • 2026.07: ย  MAR-GRPO was accepted to ACM MM 2026!
  • 2026.02: ย  Thrilled to share that our co-authored paper MaskFocus has been accepted to CVPR 2026!
  • 2026.01: ย  Excited that our collaborative work GCPO was accepted to ICLR 2026!
  • 2025.09: ย  Delighted to announce that ARSample was accepted by NeurIPS 2025!
  • 2025.06: ย  Delighted to announce that HQ-CLIP was accepted by ICCV 2025!
  • 2024.09: ย  Delighted to announce that MPI was accepted by NeurIPS 2024!
  • 2024.09: ย  I was invited to give a talk at ByteDance as the author of STAR! See slides here
  • 2024.06: ย  STAR is now available on arXiv!



๐Ÿ“ Publications

Arxiv 2024
sym

STAR: Scale-wise Text-to-image generation via Auto-Regressive representations

Xiaoxiao Ma*, Mohan Zhou*, Tao Liang, Yalong Bai, et al.

Project

  • STAR is a novel scale-wise text-to-image model that is effective and efficient in performance
  • Notably, STAR shows efficiency by requiring 2.95s to generate 512ร—512 images (compared to 6.48s for PixArt-ฮฑ)
NeurIPS 2025
sym

Towards Better & Faster Autoregressive Image Generation: From the Perspective of Entropy

Xiaoxiao Ma, Feng Zhao, Pengyang Ling, Haibo Qiu, et al.

Project

  • We revisit the sampling problem in autoregressive image generation and reveal the low and uneven information density of image tokens.
  • Based on this insight, we propose an entropy-informed decoding strategy that improves both generation quality and efficiency across diverse AR models and benchmarks.
ACM MM 2026
MAR-GRPO teaser

MAR-GRPO: Stabilized GRPO for AR-diffusion Hybrid Image Generation

Xiaoxiao Ma, Jiachen Lei, Tianfei Ren, Jie Huang, et al.

Paper Code

  • MAR-GRPO stabilizes reinforcement learning for masked autoregressive models by addressing diffusion-head-induced gradient noise.
  • Multi-trajectory expectation and uncertainty-aware token selection improve visual quality, structural consistency, and training stability.
Arxiv 2025
sym

STAGE: Stable and Generalizable GRPO for Autoregressive Image Generation

Xiaoxiao Ma, Haibo Qiu, Guohui Zhang, Zhixiong Zeng, et al.

Project

  • STAGE is the first system to study the stability and generalization of GRPO-based autoregressive visual generation.
  • Built upon Janus-Pro-7B, STAGE improves the GenEval score from 0.78 to 0.89 (โ‰ˆ14%) without compromising image quality, and its effectiveness generalizes well across benchmarks such as T2I-Compbench and ImageReward.
NeurIPS 2024
sym

Masked Pre-trained Model Enables Universal Zero-shot Denoiser

Xiaoxiao Ma*, Zhixiang Wei*, Yi Jin, Pengyang Ling, et al.

Project

  • MPI is a zero-shot denoising pipeline designed for many types of noise degradations
  • Only around 10s takes for a MPI to denoise on single noisy image
ICCV 2025
sym

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models

Zhixiang Wei*, Guangting Wang*, Xiaoxiao Ma, et al.

Project

  • A CLIP training framework trained on 1.3B bidirectional imageโ€“text pairs, combining bidirectional supervision and label classification, achieving SoTA zero-shot and retrieval performance.
CVPR 2024
sym

Stronger, Fewer, \& Superior: Harnessing Vision Foundation Models for Domain Generalized Semantic Segmentation

Zhixiang Wei*, Lin Chen*, Yi Jin*, Xiaoxiao Ma, et al.

Project

  • Rein is a PEFT framework based on vision foundation models for domain generalized semantic segmentation (DGSS) with merely 1% trainable parameters



๐Ÿ›  Tech Reports

2026.08
JoyAI-Echo 1.5 teaser

JoyAI-Echo 1.5: Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds

Core Contributor

Paper Project Code

  • A long-horizon audio-visual generation system for persistent stories and interactive worlds.



๐Ÿ’ป Experiences

  • 2026.03 - Present, JD Explore Academy, TGT Program Research Intern, Beijing.
  • 2025.04 - 2025.11, Meituan LongCat Multimodal Foundation Group, Beijing.
  • 2024.12 - 2025.03, OpenGVLab, Shanghai AI Laboratory, Shanghai.
  • 2024.04 - 2024.12, Du Xiaoman Technology, Beijing.



๐Ÿ“ Academic Service (Reviewer)

  • Conference Reviewer: AAAI (2027), ECCV (2026), ICML (2026), CVPR (2026), ICLR (2026), NeurIPS (2025, 2026)
  • Journal Reviewer: IEEE TPAMI, TMLR, IJCV



๐ŸŽ– Honors and Awards

  • 2022~2025 First Prize Scholarship of USTC for four continusous years
  • 2024 National Scholarship for Undergraduate Students



๐Ÿ“– Educations

  • 2025.09 - now, University of Science and Technology of China, Anhui, PhD candidate in Multimodal Learning
  • 2022.09 - 2025.06, University of Science and Technology of China, Anhui, Master candidate in Computer Vision
  • 2018.09 - 2022.06, China Agricultural University, Beijing. B. Eng in Computer Science