Zhe Gan

2009–2026 年に発表

167
論文数
20,364
被引用数
68
h 指数
140
i10 指数

被引用数

Zhe Gan の年別被引用数1994 年: 被引用 3 件2001 年: 被引用 1 件2002 年: 被引用 1 件2005 年: 被引用 1 件2006 年: 被引用 1 件2009 年: 被引用 1 件2010 年: 被引用 1 件2012 年: 被引用 3 件2015 年: 被引用 24 件2016 年: 被引用 40 件2017 年: 被引用 89 件2018 年: 被引用 254 件2019 年: 被引用 392 件2020 年: 被引用 737 件2021 年: 被引用 955 件2022 年: 被引用 857 件2023 年: 被引用 1,492 件2024 年: 被引用 1,627 件2025 年: 被引用 1,752 件2026 年: 被引用 586 件1995〜2000 年は被引用が無いため表示していません2003〜2004 年は被引用が無いため表示していません2007〜2008 年は被引用が無いため表示していません2011 年は被引用が無いため表示していません2013〜2014 年は被引用が無いため表示していません

引用元

国・地域

この著者を引用した国・地域の世界地図中国: 引用元論文 2,146 件、この内訳の 30.6%アメリカ合衆国: 引用元論文 1,632 件、この内訳の 23.3%イギリス: 引用元論文 444 件、この内訳の 6.3%香港: 引用元論文 231 件、この内訳の 3.3%オーストラリア: 引用元論文 219 件、この内訳の 3.1%ドイツ: 引用元論文 205 件、この内訳の 2.9%シンガポール: 引用元論文 193 件、この内訳の 2.8%カナダ: 引用元論文 188 件、この内訳の 2.7%韓国: 引用元論文 171 件、この内訳の 2.4%インド: 引用元論文 164 件、この内訳の 2.3%日本: 引用元論文 150 件、この内訳の 2.1%イタリア: 引用元論文 99 件、この内訳の 1.4%
0%30.6%その他 16.8%

分野

  • Computer Science87.2%
  • Engineering3.1%
  • Social Sciences1.6%
  • Medicine1.5%
  • Psychology1%
  • Neuroscience0.9%
  • その他4.7%

トピック

  • Multimodal Machine Learning Applications16.2%
  • Topic Modeling10.2%
  • Natural Language Processing Techniques6.8%
  • Domain Adaptation and Few-Shot Learning6.3%
  • Advanced Image and Video Retrieval Techniques5.5%
  • Human Pose and Action Recognition4%
  • その他51%

共著者

全論文

検索で開く
  1. GIT: A Generative Image-to-text Transformer for Vision and Language

    著者: , , , , , , , , - Trans. Mach. Learn. Res. 2022 被引用: 487

  2. Ferret: Refer and Ground Anything Anywhere at Any Granularity

    著者: , , , , , , , , - ICLR 2024 被引用: 569

  3. MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

    著者: , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , Alexander Toshev, Yinfei Yang - arXiv (Cornell University), CoRR 2024 被引用: 223

  4. An Empirical Study of GPT-3 for Few-Shot Knowledge-Based VQA

    著者: , , , , , , - AAAI Conference on Artificial Intelligence 2022 被引用: 278

  5. SwinBERT: End-to-End Transformers with Sparse Attention for Video Captioning

    著者: , , , , , , , - IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2022 被引用: 269

  6. Guiding Instruction-based Image Editing via Multimodal Large Language Models

    著者: , , , , , - ICLR 2024 被引用: 198

  7. Patient Knowledge Distillation for BERT Model Compression

    著者: , , , - Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), EMNLP/IJCNLP (1) 2019 被引用: 297

  8. SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

    著者: , , , , , , , - arXiv (Cornell University), CoRR 2024 被引用: 112

  9. Improve Vision Language Model Chain-of-thought Reasoning

    著者: , , , , , , , , - Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL (1) 2025 被引用: 110

  10. Multimodal Foundation Models: From Specialists to General-Purpose Assistants

    著者: , , , , , , - Foundations and Trends® in Computer Graphics and Vision, Found. Trends Comput. Graph. Vis. 2024 被引用: 134

  11. Prompting GPT-3 To Be Reliable

    著者: , , , , , , - ICLR 2023 被引用: 383

  12. Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models

    著者: , , , , , , , - NeurIPS Datasets and Benchmarks 2021 被引用: 308

  13. Scaling Up Vision-Language Pretraining for Image Captioning

    著者: , , , , , , - IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2022 被引用: 206

  14. Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

    著者: , , , , , , , , , , - arXiv (Cornell University), CoRR 2024 被引用: 87

  15. HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training

    著者: , , , , , - Conference on Empirical Methods in Natural Language Processing (EMNLP), EMNLP (1) 2020 被引用: 389

  16. VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling

    著者: , , , , , , - arXiv (Cornell University), CoRR 2021 被引用: 177

  17. GRiT: A Generative Region-to-Text Transformer for Object Understanding

    著者: , , , , , , - Lecture notes in computer science, ECCV (80) 2024 被引用: 94

  18. Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs

    著者: , , , , , , , - Lecture notes in computer science, ECCV (64) 2024 被引用: 71

  19. MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning

    著者: , , , , , , , , , , , , , , , , , , , , , , - ICLR 2025 被引用: 73

  20. FreeLB: Enhanced Adversarial Training for Natural Language Understanding

    著者: , , , , , - ICLR 2020 被引用: 521

  21. UNITER: Learning UNiversal Image-TExt Representations

    著者: , , , , , , , - Lecture notes in computer science, ECCV (30) 2020 被引用: 1,847

  22. Large-Scale Adversarial Training for Vision-and-Language Representation Learning

    著者: , , , , , - Neural Information Processing Systems, NeurIPS 2020 被引用: 559

  23. Pico-Banana-400K: A Large-Scale Dataset for Text-Guided Image Editing

    著者: , , , , , , , - ArXiv.org, CoRR 2025 被引用: 44

  24. Multimodal Autoregressive Pre-training of Large Vision Encoders

    著者: , , , , , , , , , , , , , , , - IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2025 被引用: 44