Thomas Wolf

Active 1976–2026

208
Papers
42,352
Citations
74
h-index
135
i10-index

Citations

Citations per year for Thomas Wolf1978: 2 citations1984: 1 citations1985: 2 citations1986: 3 citations1987: 3 citations1988: 1 citations1989: 5 citations1990: 2 citations1991: 7 citations1992: 13 citations1993: 7 citations1994: 14 citations1995: 11 citations1996: 15 citations1997: 34 citations1998: 17 citations1999: 39 citations2000: 51 citations2001: 38 citations2002: 25 citations2003: 40 citations2004: 35 citations2005: 43 citations2006: 44 citations2007: 30 citations2008: 31 citations2009: 35 citations2010: 29 citations2011: 31 citations2012: 39 citations2013: 52 citations2014: 71 citations2015: 60 citations2016: 63 citations2017: 100 citations2018: 148 citations2019: 480 citations2020: 1,765 citations2021: 3,662 citations2022: 3,626 citations2023: 5,016 citations2024: 4,315 citations2025: 4,527 citations2026: 1,961 citations1979–1983: no citations, so these years are not shown

Citation sources

Countries

World map of the countries and regions citing this authorUnited States: 5,828 citing papers, 24.8% of this breakdownChina: 3,470 citing papers, 14.8% of this breakdownUnited Kingdom: 1,502 citing papers, 6.4% of this breakdownGermany: 1,476 citing papers, 6.3% of this breakdownCanada: 948 citing papers, 4% of this breakdownIndia: 725 citing papers, 3.1% of this breakdownFrance: 573 citing papers, 2.4% of this breakdownItaly: 559 citing papers, 2.4% of this breakdownJapan: 552 citing papers, 2.3% of this breakdownSouth Korea: 547 citing papers, 2.3% of this breakdownAustralia: 541 citing papers, 2.3% of this breakdownSwitzerland: 458 citing papers, 2% of this breakdown
0%24.8%Other 26.9%

Fields

  • Computer Science72%
  • Medicine5.5%
  • Social Sciences4.1%
  • Engineering3.7%
  • Biochemistry, Genetics and Molecular Biology3.2%
  • Psychology2.4%
  • Other9.1%

Topics

  • Topic Modeling15.7%
  • Natural Language Processing Techniques11%
  • Multimodal Machine Learning Applications4.5%
  • Domain Adaptation and Few-Shot Learning1.9%
  • Sentiment Analysis and Opinion Mining1.9%
  • Software Engineering Research1.8%
  • Other63.2%

Coauthors

All papers

Open in search
  1. DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

    Authors: , , , - TGDK 2025 cited by 5,459

  2. Transformers: State-of-the-Art Natural Language Processing

    Authors: , , , , , , , , , , , , , , , , , , , , , - Conference on Empirical Methods in Natural Language Processing: System Demonstrations, EMNLP (Demos) 2020 cited by 8,019

  3. HuggingFace's Transformers: State-of-the-art Natural Language Processing

    Authors: , , , , , , , , , , , , , , , , , , , , , - arXiv (Cornell University), CoRR 2019 cited by 4,445

  4. StarCoder: may the source be with you!

    Authors: , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , Siva Sankalp Patel, Dmitry Abulkhanov, Marco Zocca, Manan Dey, Zhihan Zhang, Nour Fahmy, Urvashi Bhattacharyya, Wenhao Yu, Swayam Singh, Sasha Luccioni, Paulo Villegas, Maxim Kunakov, Fedor Zhdanov, Manuel Romero, Tong Lee, Nadav Timor, Ding, Jennifer, Claire Schlesinger, Hailey Schoelkopf, Jan Ebert, Tri Dao, Mayank Mishra, Alex Gu, Jennifer G. Robinson, Carolyn Jane Anderson, Brendan Dolan-Gavitt, Danish Contractor, Siva Reddy, Daniel Fried, Dzmitry Bahdanau, Yacine Jernite, Carlos Muñoz Ferrandis, Sean Hughes, Thomas Wolf, Arjun Guha, Leandro von Werra, Harm de Vries - arXiv (Cornell University), CoRR 2023 cited by 1,146

  5. Zephyr: Direct Distillation of LM Alignment

    Authors: , , , , , , , , , , , , , - arXiv (Cornell University), CoRR 2023 cited by 605

  6. Multitask Prompted Training Enables Zero-Shot Task Generalization

    Authors: , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , Jos Rozen, Abheesht Sharma, Andrea Santilli, Thibault Févry, Jason Alan Fries, Ryan Teehan, Stella Biderman, Leo Gao, Tali Bers, Thomas Wolf, Alexander M. Rush - ICLR 2022 cited by 2,035

  7. The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale

    Authors: , , , , , , , - Advances in Neural Information Processing Systems 37, NeurIPS 2024 cited by 1,056

  8. SmolLM2: When Smol Goes Big - Data-Centric Training of a Small Language Model

    Authors: , , , , , , , , , , , , , , , , , , , , , - ArXiv.org, CoRR 2025 cited by 320

  9. SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics

    Authors: , , , , , , , , , , , , , - ArXiv.org, CoRR 2025 cited by 293

  10. GAIA: a benchmark for General AI Assistants

    Authors: , , , , , - ICLR 2024 cited by 1,034

  11. The Stack: 3 TB of permissively licensed source code

    Authors: , , , , , , , , , , , , - Trans. Mach. Learn. Res. 2023 cited by 227

  12. SmolVLM: Redefining small and efficient multimodal models

    Authors: , , , , , , , , , , , , , , , , - ArXiv.org, CoRR 2025 cited by 184

  13. Datasets: A Community Library for Natural Language Processing

    Authors: , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , Alexander M. Rush, Thomas Wolf - Conference on Empirical Methods in Natural Language Processing: System Demonstrations, EMNLP (Demos) 2021 cited by 368

  14. Scaling Data-Constrained Language Models

    Authors: , , , , , , , , - J. Mach. Learn. Res. 2023 cited by 423

  15. Federated benchmarking of medical artificial intelligence with MedPerf

    Authors: , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , Victor Bittorf, Sreekar Reddy Puchala, Biagio Ricciuti, Soujanya Samineni, Eshna Sengupta, Akshay Chaudhari, Cody Coleman, Bala Desinghu, Gregory F. Diamos, Debo Dutta, Diane Feddema, Grigori Fursin, Xinyuan Huang, Satyananda Kashyap, Nicholas D. Lane, Indranil Mallick, Pietro Mascagni, Virendra Mehta, Cassiano Ferro Moraes, Vivek Natarajan, Nikola Nikolov, Nicolas Padoy, Gennady Pekhimenko, Vijay Janapa Reddi, G. Anthony Reina, Pablo Ribalta, Abhishek Singh, Jayaraman J. Thiagarajan, Jacob Albrecht, Thomas Wolf, Geralyn Miller, Huazhu Fu, Prashant Shah, Daguang Xu, Poonam Yadav, David Talby, Mark M. Awad, Jeremy P. Howard, Michael Rosenthal, Luigi Marchionni, Massimo Loda, Jason M. Johnson, Spyridon Bakas, Peter Mattson - Nature Machine Intelligence, Nat. Mac. Intell. 2023 cited by 149

  16. Transfer Learning in Natural Language Processing

    Authors: , , , - Conference of the North, NAACL-HLT (Tutorial Abstracts) 2019 cited by 579

  17. FineWeb2: One Pipeline to Scale Them All - Adapting Pre-Training Data Processing to Every Language

    Authors: , , , , , , , , , - ArXiv.org, CoRR 2025 cited by 76

  18. Grounding Large Language Models in Interactive Environments with Online Reinforcement Learning

    Authors: , , , , , - ICML 2023 cited by 272

  19. Multi-view gait recognition using 3D convolutional neural networks

    Authors: , , - IEEE International Conference on Image Processing (ICIP) 2016 cited by 246

  20. TransferTransfo: A Transfer Learning Approach for Neural Network Based Conversational Agents

    Authors: , , , - arXiv (Cornell University), CoRR 2019 cited by 301

  21. DABstep: Data Agent Benchmark for Multi-step Reasoning

    Authors: , , , , , - ArXiv.org, CoRR 2025 cited by 28

  22. antiSMASH 4.0 - improvements in chemistry prediction and gene cluster boundary identification

    Authors: , , , , , , , , , , , , , , , , , - Nucleic Acids Research, Nucleic Acids Res. 2017 cited by 1,219

  23. Movement Pruning: Adaptive Sparsity by Fine-Tuning

    Authors: , , - NeurIPS 2020 cited by 623

  24. VIMPAC: Video Pre-Training via Masked Token Prediction and Contrastive Learning

    Authors: , , , - arXiv (Cornell University), CoRR 2021 cited by 62