I am a Research Scientist at Salesforce AI Research in Palo Alto, where I work with Dr. Shafiq Joty. Language models now produce fluent answers faster than anyone can check them, so most of my work is on verification: judging a model’s reasoning step by step rather than taking its conclusions on trust.

Hard2Verify, our ACL 2026 paper, asks whether verifiers can find the faulty step in a solution to an open-ended mathematics problem instead of only grading its final answer. SFR-DeepResearch uses reinforcement learning to train a single agent that reasons and searches on its own, and a related project synthesizes training data for web agents of progressively increasing difficulty. A second thread makes mixture-of-experts models cheaper to train, most recently in Least-Loaded Expert Parallelism at ICML 2026.

I started out at BITS Pilani, Goa Campus, where I studied computer science and led the Society for Artificial Intelligence and Deep Learning. From 2021 I worked with Ameet Deshpande and Karthik Narasimhan at Princeton NLP on data augmentation for low-resource languages, and the next summer built a gesture detector for TV news at Red Hen Lab through Google Summer of Code. My undergraduate thesis, at Microsoft Research India with Dr. Sunayana Sitaram, measured what compressing a multilingual model costs in fairness; that work became an ACL 2023 paper.

At the University of Texas at Austin I completed an M.S. in Computer Science, advised by Dr. Greg Durrett and Dr. Ying Ding. My thesis research in the AI Health Lab produced MedHallu, an EMNLP 2025 benchmark for detecting medical hallucinations, and Teaching with Lies, which trains compact hallucination detectors; alongside it I worked with Dr. Durrett on repairing model-written programs and on CodeUpdateArena, and spent the summer of 2024 at Salesforce on SFR-RAG. After graduating in 2025 I returned to Salesforce — and to the question of telling correct reasoning from convincing reasoning.

Selected research

Education

  • University of Texas at Austin

    M.S. in Computer Science · advised by Greg Durrett and Ying Ding

    2023 – 2025
  • BITS Pilani, Goa Campus

    B.E. in Computer Science, minor in Data Science · thesis at Microsoft Research India

    2019 – 2023

News

  1. New preprint: Privileged Likelihood Is Not Automatically Value — token-credit checks for on-policy self-distillation.

  2. New preprint: Mixture-of-Parallelisms — a memory-efficient training stack for trillion-parameter MoE models.

  3. Hard2Verify appears at ACL 2026 as a main-conference long paper.

  4. Least-Loaded Expert Parallelism is accepted at ICML 2026.

  5. Two new preprints: Hard2Verify, a step-level verification benchmark for frontier math, and Synthesizing Agentic Data for Web Agents.

  6. Spoke at Dreamforce 2025: “How to Build Flexible Deep Research Agents” — recording on YouTube.

Show 16 earlier updatesShow fewer
  1. New technical report: SFR-DeepResearch — reinforcement learning for autonomously reasoning single agents.

  2. MedHallu is accepted at EMNLP 2025 (main conference).

  3. New preprint: EgoVLM — policy optimization for egocentric video understanding.

  4. Joined Salesforce AI Research as a Research Scientist, and graduated with an M.S. in Computer Science from UT Austin.

  5. New preprint: Teaching with Lies — curriculum DPO on synthetic negatives for hallucination detection.

  6. New preprint: MedHallu — a comprehensive benchmark for detecting medical hallucinations in LLMs.

  7. FaithEval is accepted at ICLR 2025.

  8. Technical report: SFR-RAG — towards contextually faithful LLMs.

  9. Joined Salesforce AI Research in Palo Alto as a research intern.

  10. AdaPT is accepted at NAACL 2024 (Findings).

  11. Started my M.S. in Computer Science at UT Austin.

  12. Our study on model compression and fairness is accepted at ACL 2023.

  13. Selected for Google Research Week; started a project at APPCAIR Lab in collaboration with TCS Research.

  14. Joined Microsoft Research, Bengaluru as a research intern.

  15. CIAug is accepted at NAACL 2022 and DMix at ACL 2022.

  16. Selected as a Google Summer of Code contributor with Red Hen Lab.

Publications

Google Scholar
19 entries

2026

  1. Hard2Verify: A Step-Level Verification Benchmark for Open-Ended Frontier Math

    Shrey Pandit, Austin Xu, Xuan-Phi Nguyen, Yifei Ming, Caiming Xiong, Shafiq Joty

    ACL 2026 Paper arXiv Code Dataset Verification & hallucination
  2. Least-Loaded Expert Parallelism: Load Balancing An Imbalanced Mixture-of-Experts

    Xuan-Phi Nguyen, Shrey Pandit, Austin Xu, Caiming Xiong, Shafiq Joty

    ICML 2026 arXiv Efficient training
  3. Privileged Likelihood Is Not Automatically Value: Three Checks for Token Credit in On-Policy Self-Distillation

    Xuan-Phi Nguyen, Shrey Pandit, Yiran Zhao, Anurag Koul, Zeyu Liu, Shafiq Joty

    arXiv preprint arXiv Agents & RL
  4. Mixture-of-Parallelisms: Towards Memory-Efficient Training Stack for Mixture-of-Experts Models

    Xuan-Phi Nguyen, Shrey Pandit, Yiran Zhao, Semih Yavuz, Silvio Savarese, Shafiq Joty

    arXiv preprint arXiv Efficient training
  5. Systems and Methods for Building Artificial Intelligence Agents

    Shafiq Joty, Xuan-Phi Nguyen, Shrey Pandit, et al.

    U.S. Patent Application 19/043,190 Agents & RL

2025

  1. MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models

    Shrey Pandit, Jiawei Xu, Junyuan Hong, Zhangyang Wang, Tianlong Chen, Kaidi Xu, Ying Ding

    EMNLP 2025 Paper arXiv Website Verification & hallucination
  2. FaithEval: Can Your Language Model Stay Faithful to Context, Even If “The Moon Is Made of Marshmallows”

    Yifei Ming, Senthil Purushwalkam, Shrey Pandit, Zixuan Ke, Xuan-Phi Nguyen, Caiming Xiong, Shafiq Joty

    ICLR 2025 arXiv Verification & hallucination
  3. SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents

    Xuan-Phi Nguyen, Shrey Pandit, Revanth Gangi Reddy, Austin Xu, Silvio Savarese, Caiming Xiong, Shafiq Joty

    Salesforce AI Research technical report arXiv Agents & RL
  4. Synthesizing Agentic Data for Web Agents with Progressive Difficulty Enhancement Mechanisms

    Shrey Pandit, Xuan-Phi Nguyen, Yifei Ming, Austin Xu, Jiayu Wang, Caiming Xiong, Shafiq Joty

    arXiv preprint arXiv Agents & RL
  5. Teaching with Lies: Curriculum DPO on Synthetic Negatives for Hallucination Detection

    Shrey Pandit, Ashwin Vinod, Liu Leqi, Ying Ding

    arXiv preprint arXiv Website Verification & hallucination
  6. EgoVLM: Policy Optimization for Egocentric Video Understanding

    Ashwin Vinod, Shrey Pandit, Aditya Vavre, Linshen Liu

    arXiv preprint arXiv Agents & RL

2024

  1. SFR-RAG: Towards Contextually Faithful LLMs

    Xuan-Phi Nguyen, Shrey Pandit, Senthil Purushwalkam, Austin Xu, Hailin Chen, Yifei Ming, Zixuan Ke, Silvio Savarese, Caiming Xiong, Shafiq Joty

    Salesforce AI Research technical report arXiv Blog Verification & hallucination
  2. CodeUpdateArena: Benchmarking Knowledge Editing on API Updates

    Zeyu Leo Liu, Shrey Pandit, Xi Ye, Eunsol Choi, Greg Durrett

    arXiv preprint arXiv
  3. AdaPT: A Set of Guidelines for Hyperbolic Multimodal Multilingual NLP

    Ramit Sawhney, Shrey Pandit, Vishwa Shah, Megh Thakkar, Shafiq Joty

    Findings of NAACL 2024 Paper

2023

  1. A Comparative Study on the Impact of Model Compression Techniques on Fairness in Language Models

    Krithika Ramesh, Arnav Chavan, Shrey Pandit, Sunayana Sitaram

    ACL 2023 Paper Efficient training
  2. Can LLMs Solve Generative Visual Analogies?

    Shrey Pandit, Gautam Shroff, Lovekesh Vig, Ashwin Srinivasan

    IARML Workshop @ IJCAI 2023 Paper

2022

  1. DMix: Adaptive Distance-Aware Interpolative Mixup

    Ramit Sawhney, Megh Thakkar, Shrey Pandit, Ritesh Soun, Di Jin, Diyi Yang, Lucie Flek

    ACL 2022 Paper
  2. CIAug: Equipping Interpolative Augmentation with Curriculum Learning

    Ramit Sawhney, Ritesh Soun, Shrey Pandit, Megh Thakkar, Sarvagya Malaviya, Yuval Pinter

    NAACL 2022 Paper

2020

  1. An Autoencoder-Based Approach to Simulate Sports Games

    Ashwin Vaswani, Rijul Ganguly, Het Shah, Sharan Ranjit S, Shrey Pandit, Samruddhi Bothara

    MLSA Workshop @ ECML-PKDD 2020 arXiv

Experience

  1. Salesforce AI Research, Palo Alto, CA

    May 2025 – present

    Research Scientist

    • Research on reasoning and reliability in LLMs: step-level verification of mathematical reasoning (Hard2Verify, ACL 2026), autonomous deep-research agents (SFR-DeepResearch), and agentic data synthesis for web agents.
    • Work on efficient large-scale training and inference of mixture-of-experts models (Least-Loaded Expert Parallelism, ICML 2026; Mixture-of-Parallelisms).
    • Presented “How to Build Flexible Deep Research Agents” at Dreamforce 2025, and named on U.S. Patent Application 19/043,190, “Systems and Methods for Building Artificial Intelligence Agents.”
  2. AI Health Lab, University of Texas at Austin

    Aug 2023 – May 2025

    Graduate Researcher · advised by Dr. Ying Ding and Dr. Greg Durrett

    • Introduced MedHallu (EMNLP 2025), the first benchmark designed specifically for detecting medical hallucinations in LLMs, built on 10,000 curated question–answer pairs with stratified difficulty levels.
    • Showed that hallucination detection remains difficult for semantically subtle cases, and that domain knowledge and an explicit “not sure” option substantially improve detection.
    • Developed Teaching with Lies, which trains hallucination detectors on synthetic negative examples using curriculum direct preference optimization.
    • With Dr. Greg Durrett, worked on efficient debugging of LLM-generated programs using feedback and error traces, and co-developed CodeUpdateArena, a benchmark for knowledge editing in code LLMs under synthetic API updates.
  3. Salesforce AI Research, Palo Alto, CA

    Jun 2024 – Aug 2024

    Research Intern

    • Worked with Dr. Shafiq Joty on SFR-RAG, a 9B retrieval-augmented LLM that outperformed models 10× its size on RAG benchmarks at the time of release.
    • Curated a faithfulness-focused evaluation set that tests whether a model stays grounded in the context it is given, later released as part of FaithEval (ICLR 2025).
    • Developed a “thought / observation” strategy for SFR-RAG that significantly improved multi-hop question answering.
  4. Microsoft Research, Bengaluru, India

    Jul 2022 – Jan 2023

    Research Intern

    • Worked with Dr. Sunayana Sitaram on compressing large language models with adapters, balancing computational efficiency against bias.
    • Evaluated the zero-shot performance of compressed massive multilingual models, and deployed compressed variants for downstream use.
    • Studied the effect of compression on fairness with both intrinsic and extrinsic measures (ACL 2023).
  5. Google Summer of Code, Red Hen Lab

    Jun 2022 – Sep 2022

    Contributor

    • Built an end-to-end multimodal vision transformer to detect hand gestures co-occurring with time expressions in TV news, supporting accessible captioning.
    • Framed the task as classification of body-keypoint trajectories, as part of Red Hen Lab’s work on multimodal analysis of TV news.
    • Wrote up the approach and results in a blog post.
  6. Princeton NLP

    Jun 2021 – May 2022

    Research Collaborator

    • Worked with Ameet Deshpande and Karthik Narasimhan on interpolative data augmentation (MixUp) for low-resource multilingual NLP.
    • Studied whether interpolating between training examples improves transformer models on languages with little labeled data.