I am a Research Scientist at Salesforce AI Research in Palo Alto, working with Dr. Shafiq Joty. My research centers on reasoning and reliability in large language models: step-level verification of frontier math reasoning (Hard2Verify, ACL 2026), autonomous research agents (SFR-DeepResearch), agentic data synthesis for web agents, and efficient training of mixture-of-experts models (ICML 2026).

I completed my M.S. in Computer Science at the University of Texas at Austin, advised by Dr. Greg Durrett and Dr. Ying Ding, where I worked on hallucination in LLMs — including MedHallu (EMNLP 2025), a benchmark for detecting medical hallucinations. Before that, I earned my B.E. in Computer Science from BITS Pilani, Goa Campus, with my undergraduate thesis at Microsoft Research India under Dr. Sunayana Sitaram.

Along the way I interned at Salesforce AI Research (retrieval-augmented generation, SFR-RAG) and Microsoft Research (model compression and fairness), collaborated with Princeton NLP (Ameet Deshpande, Karthik Narasimhan) on multilingual data augmentation, and contributed to Red Hen Lab through Google Summer of Code.

During my undergraduate years I served as President of the Society for Artificial Intelligence and Deep Learning (SAiDL) at BITS Goa. I am always happy to mentor students starting out in NLP research — feel free to reach out.

News

Older news
  • Feb 2025New preprint: MedHallu — a comprehensive benchmark for detecting medical hallucinations in LLMs.
  • Jan 2025FaithEval is accepted at ICLR 2025.
  • Sep 2024Technical report: SFR-RAG — towards contextually faithful LLMs.
  • Jun 2024Joined Salesforce AI Research in Palo Alto as a research intern.
  • May 2024AdaPT is accepted at NAACL 2024 (Findings).
  • Aug 2023Started my M.S. in Computer Science at UT Austin.
  • May 2023Our study on model compression and fairness is accepted at ACL 2023.
  • Jan 2023Selected for Google Research Week; started a project at APPCAIR Lab in collaboration with TCS Research.
  • Jul 2022Joined Microsoft Research, Bengaluru as a research intern.
  • May 2022CIAug is accepted at NAACL 2022 and DMix at ACL 2022.
  • May 2022Selected as a Google Summer of Code contributor with Red Hen Lab.

Publications

Also on Google Scholar.

2026

Shrey Pandit, Austin Xu, Xuan-Phi Nguyen, Yifei Ming, Caiming Xiong, Shafiq Joty
ACL 2026
Xuan-Phi Nguyen, Shrey Pandit, Austin Xu, Caiming Xiong, Shafiq Joty
ICML 2026
Xuan-Phi Nguyen, Shrey Pandit, Yiran Zhao, Anurag Koul, Zeyu Liu, Shafiq Joty
arXiv preprint
Xuan-Phi Nguyen, Shrey Pandit, Yiran Zhao, Semih Yavuz, Silvio Savarese, Shafiq Joty
arXiv preprint
Systems and Methods for Building Artificial Intelligence Agents
Shafiq Joty, Xuan-Phi Nguyen, Shrey Pandit, et al.
U.S. Patent Application 19/043,190

2025

Shrey Pandit, Jiawei Xu, Junyuan Hong, Zhangyang Wang, Tianlong Chen, Kaidi Xu, Ying Ding
EMNLP 2025
Yifei Ming, Senthil Purushwalkam, Shrey Pandit, Zixuan Ke, Xuan-Phi Nguyen, Caiming Xiong, Shafiq Joty
ICLR 2025
Xuan-Phi Nguyen, Shrey Pandit, Revanth Gangi Reddy, Austin Xu, Silvio Savarese, Caiming Xiong, Shafiq Joty
Salesforce AI Research technical report
Shrey Pandit, Xuan-Phi Nguyen, Yifei Ming, Austin Xu, Jiayu Wang, Caiming Xiong, Shafiq Joty
arXiv preprint
Shrey Pandit, Ashwin Vinod, Liu Leqi, Ying Ding
arXiv preprint
Ashwin Vinod, Shrey Pandit, Aditya Vavre, Linshen Liu
arXiv preprint

2024

Xuan-Phi Nguyen, Shrey Pandit, Senthil Purushwalkam, Austin Xu, Hailin Chen, Yifei Ming, Zixuan Ke, Silvio Savarese, Caiming Xiong, Shafiq Joty
Salesforce AI Research technical report
Zeyu Leo Liu, Shrey Pandit, Xi Ye, Eunsol Choi, Greg Durrett
arXiv preprint
Ramit Sawhney, Shrey Pandit, Vishwa Shah, Megh Thakkar, Shafiq Joty
Findings of NAACL 2024

2023

Krithika Ramesh, Arnav Chavan, Shrey Pandit, Sunayana Sitaram
ACL 2023
Shrey Pandit, Gautam Shroff, Lovekesh Vig, Ashwin Srinivasan
IARML Workshop @ IJCAI 2023

2022

Ramit Sawhney, Megh Thakkar, Shrey Pandit, Ritesh Soun, Di Jin, Diyi Yang, Lucie Flek
ACL 2022
Ramit Sawhney, Ritesh Soun, Shrey Pandit, Megh Thakkar, Sarvagya Malaviya, Yuval Pinter
NAACL 2022

2020

Ashwin Vaswani, Rijul Ganguly, Het Shah, Sharan Ranjit S, Shrey Pandit, Samruddhi Bothara
MLSA Workshop @ ECML-PKDD 2020

Experience

Salesforce AI Research, Palo Alto, CA

Research Scientist

  • Research on reasoning and reliability in LLMs: step-level verification of mathematical reasoning (Hard2Verify, ACL 2026), autonomous deep-research agents (SFR-DeepResearch), and agentic data synthesis for web agents.
  • Work on efficient large-scale training and inference of mixture-of-experts models (Least-Loaded Expert Parallelism, ICML 2026; Mixture-of-Parallelisms).

AI Health Lab, University of Texas at Austin

Graduate Research Assistant

  • Introduced MedHallu (EMNLP 2025), the first benchmark designed specifically for detecting medical hallucinations in LLMs, built on 10,000 curated question–answer pairs with stratified difficulty levels.
  • Showed that hallucination detection remains difficult for semantically subtle cases, and that domain knowledge and an explicit “not sure” option substantially improve detection.

Salesforce AI Research, Palo Alto, CA

Research Intern

  • Worked with Dr. Shafiq Joty on SFR-RAG, a 9B retrieval-augmented LLM that outperformed models 10× its size on RAG benchmarks at the time of release.
  • Curated a faithfulness-focused evaluation set (later part of FaithEval, ICLR 2025) and developed a “thought / observation” strategy that significantly improved multi-hop QA.

TAUR Lab, University of Texas at Austin

Graduate Research Assistant

  • Worked with Dr. Greg Durrett on efficient debugging of LLM-generated programs using feedback and error traces.
  • Co-developed CodeUpdateArena, a benchmark for knowledge editing in code LLMs under synthetic API updates.

Microsoft Research, Bengaluru, India

Research Intern

  • Worked with Dr. Sunayana Sitaram on compressing large language models with adapters, and on the zero-shot performance of compressed massive multilingual models.
  • Studied the effect of compression on fairness with both intrinsic and extrinsic measures (ACL 2023).

Google Summer of Code, Red Hen Lab

Contributor

  • Built an end-to-end multimodal vision transformer to detect hand gestures co-occurring with time expressions in TV news, supporting accessible captioning. Blog post.

Princeton NLP

Research Collaborator

  • Worked with Ameet Deshpande and Karthik Narasimhan on interpolative data augmentation (MixUp) for low-resource multilingual NLP.

Education

  • University of Texas at AustinM.S. in Computer Science
  • BITS Pilani, Goa CampusB.E. in Computer Science