|
Kartik Narayan
I am a 3rd year (final year) Ph.D. student in the Computer Science
department
at Johns Hopkins University, where I am a member of VIU lab, advised by Dr.Vishal Patel. My current interests
lie in multimodal LLMs, agents, long-horizon tool calling, agentic post-training, agentic search,
reasoning and video generation. My prior research focused on computer vision and its applications in
face analysis, understanding,
and recognition, with a particular emphasis on multimodal LLMs.
Currently in my Ph.D., I have been fortunate to intern at:
- Salesforce AI Research Summer 2026 (Xiangyu Peng, Prafulla Kumar Choubey, Jiaxin Zhang, Silvio Savarese, Chien-Sheng Wu): Developing efficient multimodal agentic search methods using self-distillation to provide
intermediate reward signals.
- Apple Spring 2026 (Tian Cao, Zheng Yang, Peng Zhou, Lunbo Xu, Xiaohui Tao, Zhe
Gan): Developed a multi-image, vision-centric benchmark for evaluating multimodal agentic
search, and analyzed failure modes of multimodal models in agentic search settings involving
multiple visual entities.
- Apple Summer 2025 (Navid Shiee, Peter Grasch, Chao Jia, Yinfei Yang, Zhe
Gan): Equipped multimodal LLMs with web-search tools and enabled reasoning over iterative
search calls; developed methods to incentivize self-correction and self-verification following
an initial model response.
Prior to my doctoral studies, I worked as an undergraduate researcher under
Prof. Richa Singh and
Prof. Mayank Vatsa at the
Image Analysis and Biometrics (IAB) Lab,
IIT Jodhpur, where I worked on deepfake video generation.
Email  / 
CV  / 
Google Scholar
 / 
LinkedIn  / 
Twitter  / 
Github
Open to full-time opportunities starting from December 2026.
|
|
|
Research
My current research interests are multimodal LLMs, agents, long-horizon tool calling,
agentic post-training, agentic search, reasoning and video generation. My
prior
research focused on computer vision and its applications in face analysis,
understanding, and recognition, with the goal of developing robust open-world algorithms that can be
deployed for real-world impact.
|
|
News
- [August, 2026] One Paper is accepted at IJCB 2026.
- [June, 2026] Research Intern at Salesforce AI Research for Summer 2026.
- [June, 2026] Two Papers are accepted at ECCV 2026.
-
[Feb, 2026] Research Intern at
Apple for Spring 2026.
- [January, 2026] One Paper is accepted at ICLR 2026.
- [December, 2025] One Paper is accepted at FG 2026.
- [December, 2025] One Paper is accepted at IEEE TBIOM.
- [June, 2025] One Paper is accepted at ICCV 2025.
-
[May, 2025] Research Intern at
Apple for Summer 2025.
- [April, 2025] Two Papers are accepted at FG 2025.
- [December, 2024] One Paper is accepted at AAAI 2025.
- [October, 2024] One Paper is accepted at WACV 2025.
- [October, 2024] One Paper is accepted at IEEE TBIOM.
- [January, 2024] One Paper is accepted at FG 2024.
- [August, 2023] Joined as a PhD student at VIU Lab, Johns Hopkins University.
- [May, 2023] Received CVPR Student Travel Award.
- [Feb, 2023] One Paper is Accepted at CVPR 2023.
- [September, 2022] One Paper is Accepted at IEEE Access.
- [August, 2022] One Paper is Accepted at IJCB 2022.
- [April, 2022] One Paper is Accepted at CVPRW TCV 2022.
- [January, 2022] One Paper is Accepted at IEEE SysCon 2022.
|
|
Publications
See my Google
Scholar profile for the complete and most recent publications.
|
|
|
Efficient Multimodal DeepResearch using Self-Distillation for Intermediate
Rewards
Kartik Narayan,
Becky Xiangyu Peng,
Prafulla Kumar Choubey,
Jiaxin Zhang,
Bin Lei,
Vishal M. Patel,
Silvio Savarese,
Chien-Sheng Wu
Work in Progress
|
|
|
|
MultiImage-MMSearch: Multi-Image Vision Centric Benchmark for Multimodal Agentic
Search in MLLMs
Kartik Narayan,
Tian Cao,
Zheng Yang,
Peng Zhou,
Vishal M. Patel,
Lunbo Xu,
Xiaohui Tao,
Zhe Gan
Under Review
|
|
|
|
Fork Where the Model Changes Its Mind: Belief-Shift Branching for Tree-Structured
Reinforcement Learning
Bin Lei,
Yu Li,
Prafulla Kumar Choubey,
Jiaxin Zhang,
Becky Xiangyu Peng,
Qinyuan Ye,
Kartik Narayan,
Caiwen Ding,
Silvio Savarese,
Chien-Sheng Wu
Under Review
arXiv
|
|
|
|
DeepMMSearch-R1: Empowering Multimodal LLMs in Multimodal Web Search
Kartik Narayan,
Yang Xu,
Tian Cao,
Kavya Nerella,
Vishal M. Patel,
Navid Shiee,
Peter Grasch,
Chao Jia,
Yinfei Yang,
Zhe Gan
Technical Report
arXiv
|
|
|
|
FaceMoE: Mixture of Experts for Low-Resolution Face Recognition
Kartik Narayan, Vishal M. Patel
ECCV 2026
arXiv /
project /
code
|
|
|
|
Silhouette-based Gait Foundation Model
Dingqiang Ye, Chao Fan, Kartik Narayan, Bingzhe Wu, Chengwen Luo, Jianqiang Li,
Vishal M. Patel
ECCV 2026 (Spotlight)
arXiv
|
|
|
|
RestoreVAR: Visual Autoregressive Generation for All-in-One Image Restoration
Sudarshan Rajagoplan, Kartik Narayan, Vishal M. Patel
ICLR 2026
arXiv /
project /
code
|
|
|
|
FaceXFormer: A Unified Transformer for Facial Analysis
Kartik Narayan,
Vibashan VS,
Rama Chellappa,
Vishal M. Patel
ICCV 2025
arXiv /
project /
code (341 ⭐)
|
|
|
|
SegFace: Face Segmentation of Long-Tail Classes
Kartik Narayan,
Vibashan VS,
Vishal M. Patel
AAAI 2025
arXiv /
project /
code (109 ⭐)
|
|
|
|
PETALface: Parameter Efficient Transfer Learning for Low-resolution Face
Recognition
Kartik Narayan, Nithin Gopalkrishnan Nair, Jennifer Xu, Rama Chellappa, Vishal M.
Patel
WACV 2025 (Oral)
arXiv /
project /
code
|
|
|
|
DF-Platter: Multi-subject Heterogeneous Deepfake Dataset
Kartik Narayan,
Harsh Agarwal,
Kartik Thakral,
Surbhi Mittal,
Mayank Vatsa,
Richa Singh
CVPR 2023
paper
/
poster
|
|
|
▾ Show all
publications
|
|