Csaba Szepesvári
1993–2026 年に発表
- 別表記
- Csaba Szepesvari
- 325
- 論文数
- 23,876
- 被引用数
- 74
- h 指数
- 209
- i10 指数
被引用数
引用元
国・地域
機関
分野
- Computer Science59.4%
- Decision Sciences22%
- Engineering10.4%
- Mathematics1.5%
- Physics and Astronomy1%
- Social Sciences0.8%
- その他4.9%
トピック
- Reinforcement Learning in Robotics12.9%
- Advanced Bandit Algorithms Research10.2%
- Machine Learning and Algorithms4.3%
- Artificial Intelligence in Games4.1%
- Optimization and Search Problems2.2%
- Adversarial Robustness in Machine Learning2.1%
- その他64.2%
共著者
- András György47
- Tor Lattimore26
- Yasin Abbasi-Yadkori24
- Dale Schuurmans22
- Branislav Kveton20
- Gellért Weisz16
- Rémi Munos15
- Bo Dai13
- Jincheng Mei13
- András Lörincz12
- Nevena Lazic12
- Ilja Kuzborskij11
- Mohammad Ghavamzadeh11
- Botao Hao10
- Richard S. Sutton10
- Gábor Bartók9
- Pooria Joulani9
- Zheng Wen9
- Alex Ayoub8
- Amir Massoud Farahmand8
- András Antos8
- Dávid Pál8
- Mengdi Wang8
- Shalabh Bhatnagar8
全論文
- Bandit Algorithms
著者: Tor Lattimore, Csaba Szepesvári - Cambridge University Press eBooks 2020 被引用: 851
- Bandit Based Monte-Carlo Planning
著者: Levente Kocsis, Csaba Szepesvári - Lecture notes in computer science, ECML 2006 被引用: 2,867
- To Believe or Not to Believe Your LLM
著者: Yasin Abbasi Yadkori, Ilja Kuzborskij, András György, Csaba Szepesvári - arXiv (Cornell University), CoRR 2024 被引用: 71
- Convergence Results for Single-Step On-Policy Reinforcement-Learning Algorithms
著者: Satinder Singh, Tommi S. Jaakkola, Michael L. Littman, Csaba Szepesvári - Machine Learning, Mach. Learn. 1998 被引用: 625
- Mitigating LLM Hallucinations via Conformal Abstention
著者: Yasin Abbasi Yadkori, Ilja Kuzborskij, David Stutz, András György, Adam Fisch, Arnaud Doucet, Iuliya Beloshapka, Wei-Hung Weng, Yao-Yuan Yang, Csaba Szepesvári, Ali Taylan Cemgil, Nenad Tomasev - arXiv (Cornell University), CoRR 2024 被引用: 56
- Improved Algorithms for Linear Stochastic Bandits
著者: Yasin Abbasi-Yadkori, Dávid Pál, Csaba Szepesvári - http://papers.nips.cc/paper/4417-improved-algorithms-for-linear-stochastic-bandits.pdf 2011 被引用: 2,073
- Exploration-exploitation tradeoff using variance estimates in multi-armed bandits
著者: Jean-Yves Audibert, Rémi Munos, Csaba Szepesvári - Theoretical Computer Science, Theor. Comput. Sci. 2009 被引用: 566
- Fast gradient-descent methods for temporal-difference learning with linear function approximation
著者: Richard S. Sutton, Hamid Reza Maei, Doina Precup, Shalabh Bhatnagar, David Silver, Csaba Szepesvári, Eric Wiewiora - Conference on Machine Learning, ICML 2009 被引用: 532
- Learning with a Strong Adversary
著者: Ruitong Huang, Bing Xu, Dale Schuurmans, Csaba Szepesvári - arXiv (Cornell University), CoRR 2015 被引用: 265
- Behaviour Suite for Reinforcement Learning
著者: Ian Osband, Yotam Doron, Matteo Hessel, John Aslanides, Eren Sezener, Andre Saraiva, Katrina McKinney, Tor Lattimore, Csaba Szepesvári, Satinder Singh, Benjamin Van Roy, Richard S. Sutton, David Silver, Hado van Hasselt - ICLR 2020 被引用: 212
- Online Least Squares Estimation with Self-Normalized Processes: An Application to Bandit Problems
著者: Yasin Abbasi-Yadkori, Dávid Pál, Csaba Szepesvári - arXiv (Cornell University), CoRR 2011 被引用: 61
- Empirical Bernstein stopping
著者: Volodymyr Mnih, Csaba Szepesvári, Jean-Yves Audibert - conference on Machine learning - ICML '08 2008 被引用: 188
- A Unified Analysis of Value-Function-Based Reinforcement Learning Algorithms
著者: Csaba Szepesvári, Michael L. Littman - Neural Computation, Neural Comput. 1999 被引用: 183
- Finite-Time Bounds for Fitted Value Iteration
著者: Rémi Munos, Csaba Szepesvári - http://www.sztaki.hu/~szcsaba/papers/munos08a.pdf, J. Mach. Learn. Res. 2008 被引用: 264
- Sample-Efficient Reinforcement Learning of Partially Observable Markov Games
著者: Qinghua Liu, Csaba Szepesvári, Chi Jin - Advances in Neural Information Processing Systems 35, NeurIPS 2022 被引用: 42
- Frontier LLMs Still Struggle with Simple Reasoning Tasks
著者: Alan Malek, Jiawei Ge, Nevena Lazic, Chi Jin, András György, Csaba Szepesvári - ArXiv.org, CoRR 2025 被引用: 14
- When Is Partially Observable Reinforcement Learning Not Scary?
著者: Qinghua Liu, Alan Chung, Csaba Szepesvári, Chi Jin - COLT 2022 被引用: 34
- Optimistic MLE - A Generic Model-based Algorithm for Partially Observable Sequential Decision Making
著者: Qinghua Liu, Praneeth Netrapalli, Csaba Szepesvári, Chi Jin - Symposium on Theory of Computing, STOC 2023 被引用: 24
- PAC-Bayes with Backprop
著者: Omar Rivasplata, Vikram M. Tankasali, Csaba Szepesvári - arXiv (Cornell University), CoRR 2019 被引用: 40
- Tuning Bandit Algorithms in Stochastic Environments
著者: Jean-Yves Audibert, Rémi Munos, Csaba Szepesvári - Lecture notes in computer science, ALT 2007 被引用: 161
- Multi-criteria Reinforcement Learning
著者: Zoltán Gábor, Zsolt Kalmár, Csaba Szepesvári - http://victoria.mindmaker.hu/~szepes/papers/multi-rep97.ps.gz, ICML 1998 被引用: 210
- Stochastic Low-Rank Bandits
著者: Branislav Kveton, Csaba Szepesvári, Anup Rao, Zheng Wen, Yasin Abbasi-Yadkori, S. Muthukrishnan - arXiv (Cornell University), CoRR 2017 被引用: 33
- The Asymptotic Convergence-Rate of Q-learning
著者: Csaba Szepesvári - http://www.ualberta.ca/~szepesva/papers/nips97.ps.pdf 1997 被引用: 183
- On Multi-objective Policy Optimization as a Tool for Reinforcement Learning
著者: Abbas Abdolmaleki, Sandy H. Huang, Giulia Vezzani, Bobak Shahriari, Jost Tobias Springenberg, Shruti Mishra, Dhruva TB, Arunkumar Byravan, Konstantinos Bousmalis, András György, Csaba Szepesvári, Raia Hadsell, Nicolas Heess, Martin A. Riedmiller - arXiv (Cornell University), CoRR 2021 被引用: 22
