Tong Lu

Active 1992–2026

Also published as
Tong Lü
526
Papers
35,130
Citations
78
h-index
291
i10-index

Citations

Citations per year for Tong Lu1968: 1 citations1988: 2 citations1994: 5 citations1997: 2 citations1998: 1 citations1999: 5 citations2000: 3 citations2001: 11 citations2002: 25 citations2003: 17 citations2004: 44 citations2005: 73 citations2006: 87 citations2007: 76 citations2008: 70 citations2009: 81 citations2010: 83 citations2011: 76 citations2012: 61 citations2013: 65 citations2014: 106 citations2015: 94 citations2016: 154 citations2017: 182 citations2018: 239 citations2019: 413 citations2020: 680 citations2021: 1,008 citations2022: 1,580 citations2023: 2,482 citations2024: 3,769 citations2025: 5,739 citations2026: 2,868 citations2027: 3 citations1969–1987: no citations, so these years are not shown1989–1993: no citations, so these years are not shown1995–1996: no citations, so these years are not shown

Citation sources

Countries

World map of the countries and regions citing this authorChina: 7,464 citing papers, 42.9% of this breakdownUnited States: 2,359 citing papers, 13.6% of this breakdownUnited Kingdom: 661 citing papers, 3.8% of this breakdownIndia: 595 citing papers, 3.4% of this breakdownHong Kong: 535 citing papers, 3.1% of this breakdownGermany: 508 citing papers, 2.9% of this breakdownSouth Korea: 457 citing papers, 2.6% of this breakdownAustralia: 439 citing papers, 2.5% of this breakdownSingapore: 342 citing papers, 2% of this breakdownCanada: 338 citing papers, 2% of this breakdownJapan: 300 citing papers, 1.7% of this breakdownFrance: 268 citing papers, 1.5% of this breakdown
0%42.9%Other 18%

Fields

  • Computer Science60.7%
  • Engineering12.9%
  • Medicine11.2%
  • Biochemistry, Genetics and Molecular Biology5.7%
  • Neuroscience1.9%
  • Environmental Science1.3%
  • Other6.3%

Topics

  • Multimodal Machine Learning Applications7.4%
  • Advanced Neural Network Applications6%
  • Advanced Image and Video Retrieval Techniques3.2%
  • Domain Adaptation and Few-Shot Learning3.1%
  • Topic Modeling2.5%
  • Handwritten Text Recognition Techniques2.5%
  • Other75.3%

Coauthors

All papers

Open in search
  1. Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions

    Authors: , , , , , , , , - IEEE/CVF International Conference on Computer Vision (ICCV) 2021 cited by 4,694

  2. InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

    Authors: , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , Wenqi Shao, Junjun He, Yingtong Xiong, Wenwen Qu, Peng Sun, Penglong Jiao, Han Lv, Lijun Wu, Kaipeng Zhang, Huipeng Deng, Jiaye Ge, Kai Chen, Limin Wang, Min Dou, Lewei Lu, Xizhou Zhu, Tong Lu, Dahua Lin, Yu Qiao, Jifeng Dai, Wenhai Wang - ArXiv.org, CoRR 2025 cited by 1,446

  3. Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

    Authors: , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , Jiaye Ge, Kai Chen, Zhang, Kaipeng, Wang, Limin, Min Dou, Lewei Lu, Zhu, Xizhou, Tong Lü, Dahua Lin, Yu Qiao, Jifeng Dai, Wenhai Wang - arXiv (Cornell University), CoRR 2024 cited by 1,368

  4. PVT v2: Improved baselines with Pyramid Vision Transformer

    Authors: , , , , , , , , - Computational Visual Media, Comput. Vis. Media 2022 cited by 2,178

  5. InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

    Authors: , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , Tianyi Zhang, Songze Li, Xiangyu Zhao, Haodong Duan, Nianchen Deng, Bin Fu, Yinan He, Yi Wang, Conghui He, Botian Shi, Junjun He, Yingtong Xiong, Han Lv, Lijun Wu, Wenqi Shao, Kaipeng Zhang, Huipeng Deng, Biqing Qi, Jiaye Ge, Qipeng Guo, Wenwei Zhang, Songyang Zhang, Maosong Cao, Junyao Lin, Kexian Tang, Jianfei Gao, Haian Huang, Yuzhe Gu, Chengqi Lyu, Huanze Tang, Rui Wang, Haijun Lv, Wanli Ouyang, Limin Wang, Min Dou, Xizhou Zhu, Tong Lu, Dahua Lin, Jifeng Dai, Weijie Su, Bowen Zhou, Kai Chen, Yu Qiao, Wenhai Wang, Gen Luo - ArXiv.org, CoRR 2025 cited by 1,013

  6. How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

    Authors: , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , Tong Lu, Dahua Lin, Yu Qiao, Jifeng Dai, Wenhai Wang, Wenhai Wang, Wenhai Wang - Science China Information Sciences, Sci. China Inf. Sci. 2024 cited by 701

  7. InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

    Authors: , , , , , , , , , , , , , , - arXiv (Cornell University), CoRR 2023 cited by 522

  8. Intern VL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

    Authors: , , , , , , , , , , , , , , - IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2024 cited by 414

  9. Vision Transformer Adapter for Dense Predictions

    Authors: , , , , , , - ICLR 2023 cited by 892

  10. VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

    Authors: , , , , , , , , , , - NeurIPS 2023 cited by 699

  11. Shape Robust Text Detection With Progressive Scale Expansion Network

    Authors: , , , , , , - IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2019 cited by 844

  12. BEVFormer: Learning Bird's-Eye-View Representation from Multi-Camera Images via Spatiotemporal Transformers

    Authors: , , , , , , , - Lecture notes in computer science, ECCV (9) 2022 cited by 1,187

  13. BEVFormer: Learning Bird's-Eye-View Representation From LiDAR-Camera via Spatiotemporal Transformers

    Authors: , , , , , , , - IEEE Transactions on Pattern Analysis and Machine Intelligence, IEEE Trans. Pattern Anal. Mach. Intell. 2024 cited by 174

  14. Video Mamba Suite: State Space Model as a Versatile Alternative for Video Understanding

    Authors: , , , , , , , , , - arXiv (Cornell University), CoRR 2024 cited by 103

  15. VideoLLM: Modeling Video Sequence with Large Language Models

    Authors: , , , , , , , , , , - arXiv (Cornell University), CoRR 2023 cited by 116

  16. Neutrophil-induced ferroptosis promotes tumor necrosis in glioblastoma progression

    Authors: , , , , , , , , , , , , , , , , - Nature Communications 2020 cited by 411

  17. LLDiffusion: Learning degradation representations in diffusion models for low-light image enhancement

    Authors: , , , , , , , - Pattern Recognition, Pattern Recognit. 2025 cited by 100

  18. Vision-RWKV: Efficient and Scalable Visual Perception with RWKV-Like Architectures

    Authors: , , , , , , , , , - ICLR 2025 cited by 155

  19. Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models

    Authors: , , , , , , , , , , , , , , , , , , , , , , , , , , - ArXiv.org, CoRR 2025 cited by 73

  20. The All-Seeing Project: Towards Panoptic Visual Recognition and Understanding of the Open World

    Authors: , , , , , , , , , , , , , - ICLR 2024 cited by 130

  21. InternImage: Exploring Large-Scale Vision Foundation Models with Deformable Convolutions

    Authors: , , , , , , , , , , , - IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2023 cited by 898

  22. MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer

    Authors: , , , , , , , , , , , , - arXiv (Cornell University), CoRR 2024 cited by 67

  23. VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks

    Authors: , , , , , , , , , , , , - Advances in Neural Information Processing Systems 37, NeurIPS 2024 cited by 172

  24. OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text

    Authors: , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , Junjun He, Zhongying Tu, Tong Lu, Yali Wang, Limin Wang, Dahua Lin, Yu Qiao, Botian Shi, Conghui He, Jifeng Dai - ICLR 2025 cited by 58