I am a Research Scientist at Salesforce AI Research in Palo Alto, working with Dr. Shafiq Joty. My research centers on reasoning and reliability in large language models: step-level verification of frontier math reasoning (Hard2Verify, ACL 2026), autonomous research agents (SFR-DeepResearch), agentic data synthesis for web agents, and efficient training of mixture-of-experts models (ICML 2026).
I completed my M.S. in Computer Science at the University of Texas at Austin, advised by Dr. Greg Durrett and Dr. Ying Ding, where I worked on hallucination in LLMs — including MedHallu (EMNLP 2025), a benchmark for detecting medical hallucinations. Before that, I earned my B.E. in Computer Science from BITS Pilani, Goa Campus, with my undergraduate thesis at Microsoft Research India under Dr. Sunayana Sitaram.
Along the way I interned at Salesforce AI Research (retrieval-augmented generation, SFR-RAG) and Microsoft Research (model compression and fairness), collaborated with Princeton NLP (Ameet Deshpande, Karthik Narasimhan) on multilingual data augmentation, and contributed to Red Hen Lab through Google Summer of Code.
During my undergraduate years I served as President of the Society for Artificial Intelligence and Deep Learning (SAiDL) at BITS Goa. I am always happy to mentor students starting out in NLP research — feel free to reach out.
Publications
Also on Google Scholar.
2026
Shrey Pandit, Austin Xu, Xuan-Phi Nguyen, Yifei Ming, Caiming Xiong, Shafiq Joty
ACL 2026
Xuan-Phi Nguyen, Shrey Pandit, Austin Xu, Caiming Xiong, Shafiq Joty
ICML 2026
Xuan-Phi Nguyen, Shrey Pandit, Yiran Zhao, Anurag Koul, Zeyu Liu, Shafiq Joty
arXiv preprint
Xuan-Phi Nguyen, Shrey Pandit, Yiran Zhao, Semih Yavuz, Silvio Savarese, Shafiq Joty
arXiv preprint
Systems and Methods for Building Artificial Intelligence Agents
Shafiq Joty, Xuan-Phi Nguyen, Shrey Pandit, et al.
U.S. Patent Application 19/043,190
2025
Shrey Pandit, Jiawei Xu, Junyuan Hong, Zhangyang Wang, Tianlong Chen, Kaidi Xu, Ying Ding
EMNLP 2025
Yifei Ming, Senthil Purushwalkam, Shrey Pandit, Zixuan Ke, Xuan-Phi Nguyen, Caiming Xiong, Shafiq Joty
ICLR 2025
Xuan-Phi Nguyen, Shrey Pandit, Revanth Gangi Reddy, Austin Xu, Silvio Savarese, Caiming Xiong, Shafiq Joty
Salesforce AI Research technical report
Shrey Pandit, Xuan-Phi Nguyen, Yifei Ming, Austin Xu, Jiayu Wang, Caiming Xiong, Shafiq Joty
arXiv preprint
Shrey Pandit, Ashwin Vinod, Liu Leqi, Ying Ding
arXiv preprint
Ashwin Vinod, Shrey Pandit, Aditya Vavre, Linshen Liu
arXiv preprint
2024
Xuan-Phi Nguyen, Shrey Pandit, Senthil Purushwalkam, Austin Xu, Hailin Chen, Yifei Ming, Zixuan Ke, Silvio Savarese, Caiming Xiong, Shafiq Joty
Salesforce AI Research technical report
Zeyu Leo Liu, Shrey Pandit, Xi Ye, Eunsol Choi, Greg Durrett
arXiv preprint
Ramit Sawhney, Shrey Pandit, Vishwa Shah, Megh Thakkar, Shafiq Joty
Findings of NAACL 2024
2023
Krithika Ramesh, Arnav Chavan, Shrey Pandit, Sunayana Sitaram
ACL 2023
Shrey Pandit, Gautam Shroff, Lovekesh Vig, Ashwin Srinivasan
IARML Workshop @ IJCAI 2023
2022
Ramit Sawhney, Megh Thakkar, Shrey Pandit, Ritesh Soun, Di Jin, Diyi Yang, Lucie Flek
ACL 2022
Ramit Sawhney, Ritesh Soun, Shrey Pandit, Megh Thakkar, Sarvagya Malaviya, Yuval Pinter
NAACL 2022
2020
Ashwin Vaswani, Rijul Ganguly, Het Shah, Sharan Ranjit S, Shrey Pandit, Samruddhi Bothara
MLSA Workshop @ ECML-PKDD 2020
Experience
Salesforce AI Research, Palo Alto, CA
May 2025 – present
Research Scientist
- Research on reasoning and reliability in LLMs: step-level verification of mathematical reasoning (Hard2Verify, ACL 2026), autonomous deep-research agents (SFR-DeepResearch), and agentic data synthesis for web agents.
- Work on efficient large-scale training and inference of mixture-of-experts models (Least-Loaded Expert Parallelism, ICML 2026; Mixture-of-Parallelisms).
AI Health Lab, University of Texas at Austin
Jun 2024 – May 2025
Graduate Research Assistant
- Introduced MedHallu (EMNLP 2025), the first benchmark designed specifically for detecting medical hallucinations in LLMs, built on 10,000 curated question–answer pairs with stratified difficulty levels.
- Showed that hallucination detection remains difficult for semantically subtle cases, and that domain knowledge and an explicit “not sure” option substantially improve detection.
Salesforce AI Research, Palo Alto, CA
Jun 2024 – Aug 2024
Research Intern
- Worked with Dr. Shafiq Joty on SFR-RAG, a 9B retrieval-augmented LLM that outperformed models 10× its size on RAG benchmarks at the time of release.
- Curated a faithfulness-focused evaluation set (later part of FaithEval, ICLR 2025) and developed a “thought / observation” strategy that significantly improved multi-hop QA.
TAUR Lab, University of Texas at Austin
Jul 2023 – Jun 2024
Graduate Research Assistant
- Worked with Dr. Greg Durrett on efficient debugging of LLM-generated programs using feedback and error traces.
- Co-developed CodeUpdateArena, a benchmark for knowledge editing in code LLMs under synthetic API updates.
Microsoft Research, Bengaluru, India
Jul 2022 – Jan 2023
Research Intern
- Worked with Dr. Sunayana Sitaram on compressing large language models with adapters, and on the zero-shot performance of compressed massive multilingual models.
- Studied the effect of compression on fairness with both intrinsic and extrinsic measures (ACL 2023).
Google Summer of Code, Red Hen Lab
Jun 2022 – Sep 2022
Contributor
- Built an end-to-end multimodal vision transformer to detect hand gestures co-occurring with time expressions in TV news, supporting accessible captioning. Blog post.
Princeton NLP
Jun 2021 – May 2022
Research Collaborator
- Worked with Ameet Deshpande and Karthik Narasimhan on interpolative data augmentation (MixUp) for low-resource multilingual NLP.