Shiyu Xuan

Shiyu Xuan

Associate Professor

School of Computer Science and Engineering
Nanjing University of Science and Technology, China

About Me

I am now an Associate Professor in School of Computer Science and Engineering of Nanjing University of Science and Technology, where I am a member of IMAG. I got my Ph.D. degree at Peking University under the supervision of Prof. Shiliang Zhang. Before that, I received a master's degree from the University of Chinese Academy of Sciences in 2020 and received a bachelor’s degree in School of Electronic Information and Communications at Huazhong University of Science and Technology (HUST) in 2017. I am a recipient of the ACM SIGMM China Outstanding Doctoral Dissertation Award and a member of the CCF Multimedia Technical Committee Secretariat. My research interests include computer vision and deep learning, focusing on multimodal learning, unsupervised learning, object tracking, domain adaptation, and domain generalization.

News

  • [2026. 09] One paper on unified multi-modal object tracking was accepted by T-PAMI.
  • [2026. 04] One paper on video temporal grounding was accepted by T-MM. Congratulations to Hao Zhang on his first paper at NJUST!
  • [2026. 01] One paper on zero-shot HOI detection was accepted by ICLR.
  • [2025. 12] One paper on spike camera optical flow estimation was accepted by T-PAMI. Congratulations to our collaborator Rui Zhao!
  • [2025. 09] I received the ACM SIGMM China Outstanding Doctoral Dissertation Award.
  • [2025. 09] I received the NSFC Young Scientists Fund (Category C) and the Jiangsu Provincial Natural Science Fund for Young Scholars.
  • [2024. 09] I got 2024 CCF-Baidu Open Fund. Thanks for all.
  • [2024. 09] One paper related to multi-modal prompt learning was accepted by T-IP.
  • [2024. 05] One paper related to incremental learning was accepted by IJCV.
  • [2024. 05] I successfully defended my Ph.D. Thesis.
  • [2024. 03] One paper on multi-modal large language model was accepted by CVPR.
  • [2024. 03] My co-author's paper on human pose estimation was accepted by CVPR.
  • [2023. 12] After a tough time, one paper on long-tailed recognition was accepted by AAAI.
  • [2023. 12] My co-author's paper on long-tailed recognition was accepted by AAAI.
  • [2022. 06] I got the Peking University President Scholarship.
  • [2022. 05] The code of IIDS was released.
  • [2022. 03] The extended version of IICS was accepted by T-PAMI.
  • [2021. 06] I was invited to give a tutorial on person ReID on ICME2021.
  • [2021. 03] One paper on unsupervised person ReID was accepted by CVPR.
  • [2020. 10] One paper on long-term object tracking was accepted by Pattern Recognition.
  • [2020. 09] I started pursuing Ph.D at Peking University.
  • [2020. 06] I started an internship in SenseTime.
  • [2020. 06] I graduated from Technology and Engineering Center for Space Utilization, Chinese Academy of Sciences.
  • [2019. 07] I achieved 5th place (Overall) and 2nd place (Long-term tracking) in Single Object Tracking, VisDrone Challenge, ICCVW.
  • [2019. 05] I achieved 1st place in test set of OxUvA long-term tracking dataset.
  • [2019. 04] One paper on object tracking in satellite video was accepted by TGRS.

Publications

  • Xuan, S., Z. Li, J. Tang, & M. Zhao. Diff-MM: Exploring Pre-Trained Text-to-Image Generation Models for Unified Multi-Modal Object Tracking. IEEE Transactions on Pattern Analysis and Machine Intelligence (2026). [Paper], [Code]
  • H. Zhang, S. Xuan*, & Z. Li. Textual and Temporal-Guided Feature Decoupling for Video Temporal Grounding. IEEE Transactions on Multimedia (2026). [Paper]
  • Xuan, S., D. Wang, Z. Li, & J. Tang. Zero-Shot HOI Detection with MLLM-Based Detector-Agnostic Interaction Recognition. The Fourteenth International Conference on Learning Representations (2026). [Paper], [Code]
  • Y. Zhang, S. Xuan, & Z. Li. Robust Object Detection in Adverse Weather with Feature Decorrelation via Independence Learning. Pattern Recognition (2026). [Paper]
  • F. Xu, L. Jin, Y. Sun, S. Xuan, & Z. Li. Dual-Estimator: Decoupling Global and Local Semantic Shift for Drift Compensation in Class-Incremental Learning. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (2026). [Paper]
  • F. Xu, L. Jin, Y. Sun, S. Xuan, & Z. Li. Class-Aware Drift Compensation for Non-Uniform Semantic Shift in Continual Learning. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (2026). [Paper]
  • Y. Zhang, S. Xuan*, & Z. Li. Object Detection under Low-Light Conditions via Degradation Learning Driven by Foundation Models. ACM Transactions on Multimedia Computing, Communications and Applications (2026). [Paper]
  • R. Zhao, R. Xiong, D. Wang, S. Xuan, J. Zhang, X. Fan, & T. Huang. Spike Camera Optical Flow Estimation Based on Continuous Spike Streams. IEEE Transactions on Pattern Analysis and Machine Intelligence (2025). [Paper]
  • D. Wang, J. Duan, L. Wen, S. Xuan, H. Chen, & S. Zhang. Generalizable Object Keypoint Localization from Generative Priors. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (2025). [Paper]
  • Xuan, S., Yang M, & Zhang, S. Incremental Model Enhancement via Memory-based Contrastive Learning. International Journal of Computer Vision (2025). [Paper]
  • Xuan, S., Yang M, & Zhang, S. Adapting Vision-Language Models via Learning to Inject Knowledge. IEEE Transactions on Image Processing (2024). [Paper]
  • Xuan, S., Guo, Q, Yang M, & Zhang, S. Pink: Unveiling the Power of Referential Comprehension for Multi-modal LLMs. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (2024). [Paper], [Code]
  • Xuan, S., & Zhang, S. Decoupled Contrastive Learning for Long-Tailed Recognition. Proceedings of the AAAI Conference on Artificial Intelligence (2024). [Paper], [Code]
  • Wang, D., Xuan, S., & Zhang, S. LocLLM: Exploiting Generalizable Human Keypoint Localization via Large Language Model. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (2024). [Paper], [Code]
  • Cong, C., Xuan, S., Liu, S., Zhang, S., Pagnucco, M., & Song, Y. Decoupled Optimisation for Long-tailed Visual Recognition. Proceedings of the AAAI Conference on Artificial Intelligence (2024). [Paper]
  • Xuan, S., & Zhang, S. Intra-Inter Domain Similarity for Unsupervised Person Re-Identification. IEEE Transactions on Pattern Analysis and Machine Intelligence (2022). [Paper], [Code]
  • Zhao, M., Li, S., Xuan, S., Kou, L., Gong, S., & Zhou, Z. SatSOT: A Benchmark Dataset for Satellite Video Single Object Tracking. IEEE Transactions on Geoscience and Remote Sensing (2022). [Paper]
  • Xuan, S., & Zhang, S. Intra-Inter Camera Similarity for Unsupervised Person Re-Identification. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (2021). [Paper], [Code]
  • Xuan S, Li S, Zhao Z, et al. Siamese Networks with Distractor-Reduction Method for Long-Term Visual Object Tracking. Pattern Recognition (2021). [Paper]
  • Xuan, S, Li S, Zhao Z, et al. Rotation Adaptive Correlation Filter for Moving Object Tracking in Satellite Videos. Neurocomputing (2021). [Paper]
  • Xuan S, Li S, Han M, et al. Object Tracking in Satellite Videos by Improved Correlation Filters with Motion Estimations. IEEE Transactions on Geoscience and Remote Sensing (2019). [Paper], [Code]
  • Du, D., Zhu, P., Wen, L., Bian, X., Ling, H., Hu, Q., Zheng, J., Peng, T., Wang, X., et al. VisDrone-SOT2019: The Vision Meets Drone Single Object Tracking Challenge Results. IEEE/CVF International Conference on Computer Vision Workshop (2019). [Paper]

* Corresponding author

Professional Activities

  • Journal Reviewer of TPAMI, IJCV, TMM, ISPRS, TIP, TNNLS, TCSVT, TGRS.
  • Conference Reviewer of ICCV, ECCV, CVPR, NeurIPS, ICLR, AAAI, ACM MM.